In this episode of the Verisian Community Podcast, we're thrilled to welcome Mansi Thakrar, a Statistical Programmer at Massachusetts General Hospital. This marks our second installment dedicated to discussing good programming practices, where Mansi shares her insights into what constitutes good statistical code in biostatistics. Additionally, we delve into the challenges and opportunities presented by AI in clinical programming, and Mansi provides practical tips tailored for new statistical programmers.
Guest: Mansi Thakrar, statistical programmer at Massachusetts General Hospital working on the NIH-funded RECOVER project studying long-term post-COVID symptoms across tens of thousands of patients; previously completed a summer internship at Pfizer.
Host: Tomás Sabat Stöfsel, CEO & Co-Founder, Verisian
Topics: Defining good clinical programming practice, self-documenting and shareable code, onboarding onto a new team's coding conventions, robustness and CDISC standards (SDTM, ADaM), regulatory compliance considerations, automating recurring reports, using AI as a debugging and learning assistant rather than a code generator, advice for early-career programmers
The full episode can be found on YouTube, Apple Podcasts or Spotify.
Mansi's definition of good clinical programming centers on three qualities: functionality (each component works independently), shareability (someone else can pick it up, understand it, and modify it), and self-documentation, code that explains what it's doing without requiring a separate manual. She credits more experienced colleagues on her team with habits she's since adopted herself: sectioning and indexing long scripts (hundreds of lines) so specific functions are easy to locate, and converting reusable logic into custom functions or macros so a colleague doesn't have to recreate work that's already been solved once. Early in her own career, she admits she focused purely on meeting deadlines and didn't consider whether her code would be useful to anyone else, a mindset she's since reversed.
Mansi describes the first few months on any new team as inherently disorienting, since every team develops its own conventions, and recommends building a personal cheat sheet from day one: noting recurring custom functions, commonly used packages (her team relies heavily on dplyr and tidyr), and coding patterns as you encounter them, rather than expecting to absorb everything at once. She encourages asking direct questions during the interview process itself, how the team codes, what's expected, requesting sample code, and pushes back on the instinct to stay quiet out of fear of looking inexperienced. Her manager's central challenge, she notes, is standardizing how a team with wildly different experience levels (from ten years in R to ten months) all code similarly enough that scripts remain genuinely shareable across the group.
Mansi's team has automated its recurring reporting workflow to the point where scripts for quarterly, weekly, and daily enrollment reports run unattended at 4am and email results directly to her and colleagues, requiring her only to forward the output to the appropriate regulatory office. Selecting and running the relevant script for a given patient cohort (adult, pregnancy, pediatric) produces around ten CSV reports with no manual coding required unless a quarter's protocol changes. Her guiding principle, drawing an explicit parallel to letting money make money, is to design code so it doesn't need to be rewritten every time a report is due.
Mansi is direct that she does not use AI to write code for her, but relies on it heavily as a debugging and learning aid, pasting error messages into ChatGPT the same way she'd search Stack Overflow, a habit she traces back to a graduate-school professor's advice that a coder's core skill is knowing what to Google. She notes SAS's own error messages are often more explanatory and directly actionable than R's, which she finds comparatively unhelpful when something breaks. She draws an explicit analogy to Grammarly: a tool for catching and explaining mistakes, not one that writes the content for you, and warns that outsourcing actual coding to AI is a step too far ("just give them your salary in that case"). Recurring AI-flagged mistakes get logged into her personal cheat sheet so she stops needing to ask AI or search for the same fix twice. Separately, she and Tomás both note that despite SAS having comparatively little public training data available relative to Python, LLMs are still surprisingly accurate at producing SAS code.
Asked about regulatory expectations, Mansi notes compliance requirements vary by organization and data sensitivity, PHI and FDA-facing datasets carry more explicit requirements, and recommends early-career programmers proactively study CDISC standards (SDTM, ADaM) even if not deeply covered in graduate coursework; she supplemented her own education with certificate courses covering how to transform raw Excel or CSV data into CDISC-standard formats. Her broader advice is to keep updating that knowledge continuously, since the standards themselves keep evolving, and that the learning curve flattens considerably after hands-on practice (she cites having personally transformed 10 to 11 databases into standard formats as the point things "stopped feeling difficult").
"The code should tell a story of its own. The code should explain itself and it should be able to present what the author has put on the efforts on." — Mansi Thakrar
"It's very systematic in the way that if your code has to be run over and over again, make it as automated as possible... let your code do your work for you." — Mansi Thakrar
"I would say rather than using AI to do your job for you, use AI smartly where you can take it as assistance." — Mansi Thakrar
"If I keep making the same mistake again and again, I just take that and paste it in my cheat sheet... that helps me a lot to make sure that I'm not making the same mistakes again." — Mansi Thakrar
"Be ready for a lot of change. Be flexible. Try to learn as much as you can. No matter how much you do learn, your job is going to teach you 10 new things that you didn't know." — Mansi Thakrar
What does Mansi consider good clinical programming practice?
Code that is functional, shareable, and self-documenting, meaning it explains its own purpose through structure (sections, indexing, comments) and reusable functions, so a colleague can understand, modify, or extend it without tracking down the original author.
How did Mansi's team automate its recurring clinical reports?
Scripts for quarterly, weekly, and daily enrollment reports run unattended overnight (around 4am), automatically producing CSV outputs and emailing them to the team, reducing what used to be manual reporting work to forwarding the output to the relevant regulatory office.
How does Mansi use AI tools like ChatGPT in her programming work?
Strictly as a debugging and learning assistant, pasting error messages to understand what went wrong, similar to searching Stack Overflow, rather than asking AI to generate code outright. She logs recurring fixes into a personal cheat sheet to reduce repeat lookups over time.
Why does Mansi say SAS is sometimes easier to debug than R?
She finds SAS's error messages more explanatory and actionable, often suggesting what the correct code should have been, whereas R's error messages, in her experience, explain the underlying issue less clearly.
What advice does Mansi give to programmers just starting out or joining a new team?
Expect the first three to six months to be disorienting as you learn a team's specific coding conventions; build a personal cheat sheet of recurring patterns and functions, ask direct questions without hesitation, and stay flexible, since every job teaches new things regardless of prior experience.
The Verisian Community Podcast brings together experts in clinical trials to exchange innovative ideas and best practices central to clinical reporting, submission and review. Aligned with Verisian's mission to accelerate the evaluation and market launch of new medical treatments, each episode features expert insights, with guests ranging from statistical programmers to medical writers, to discuss the challenges and opportunities of the latest software and technology.
You can listen to us on YouTube, Apple Podcasts or Spotify.
