We use cookies to ensure that we give you the best experience on our website.
‍Read about cookies preferences.
by
Tomás Sabat Stöfsel
February 5, 2024
1 min read
Good Programming Practice in Clinical Programming II

‍

Episode Summary

In this episode of the Verisian Community Podcast, we're thrilled to welcome Mansi Thakrar, a Statistical Programmer at Massachusetts General Hospital. This marks our second installment dedicated to discussing good programming practices, where Mansi shares her insights into what constitutes good statistical code in biostatistics. Additionally, we delve into the challenges and opportunities presented by AI in clinical programming, and Mansi provides practical tips tailored for new statistical programmers.

Key Takeaways

  • Mansi Thakrar defines good clinical code around functionality, shareability, and self-documentation: code should "tell a story of its own," working independently while remaining easy for someone else to understand, modify, and extend without needing to track down the original author.
  • Her team automated recurring quarterly, monthly, and daily enrollment reports so thoroughly that scripts run unattended overnight (4am) and email results directly, turning what used to be manual reporting work into a "select all and run" process, her core philosophy being to "let your code do your work for you."
  • She uses AI (ChatGPT) specifically as a debugging assistant, pasting error messages to get an explanation, rather than to generate code outright, and maintains a personal cheat sheet of recurring syntax mistakes and functions so she increasingly doesn't need to ask AI or search the internet for repeat issues.
  • She notes SAS's error messages are often more explanatory and actionable than R's, and separately observes that despite SAS having far less public training data available than Python, LLMs are still surprisingly accurate at producing SAS code.
  • New team members should expect a difficult onboarding period (Mansi describes her own transition from SAS to R at Massachusetts General Hospital as an "imposter syndrome" experience), and her practical advice is to build a personal cheat sheet of team-specific coding patterns and packages, ask questions freely, and expect the first several months to be dedicated to learning the team's system.

Episode Info

Guest: Mansi Thakrar, statistical programmer at Massachusetts General Hospital working on the NIH-funded RECOVER project studying long-term post-COVID symptoms across tens of thousands of patients; previously completed a summer internship at Pfizer.

Host: Tomás Sabat Stöfsel, CEO & Co-Founder, Verisian

Topics: Defining good clinical programming practice, self-documenting and shareable code, onboarding onto a new team's coding conventions, robustness and CDISC standards (SDTM, ADaM), regulatory compliance considerations, automating recurring reports, using AI as a debugging and learning assistant rather than a code generator, advice for early-career programmers

The full episode can be found on YouTube, Apple Podcasts or Spotify.

Good Code "Tells a Story of Its Own"

Mansi's definition of good clinical programming centers on three qualities: functionality (each component works independently), shareability (someone else can pick it up, understand it, and modify it), and self-documentation, code that explains what it's doing without requiring a separate manual. She credits more experienced colleagues on her team with habits she's since adopted herself: sectioning and indexing long scripts (hundreds of lines) so specific functions are easy to locate, and converting reusable logic into custom functions or macros so a colleague doesn't have to recreate work that's already been solved once. Early in her own career, she admits she focused purely on meeting deadlines and didn't consider whether her code would be useful to anyone else, a mindset she's since reversed.

Onboarding to a New Team's System, and Why the First Three Months Feel Miserable

Mansi describes the first few months on any new team as inherently disorienting, since every team develops its own conventions, and recommends building a personal cheat sheet from day one: noting recurring custom functions, commonly used packages (her team relies heavily on dplyr and tidyr), and coding patterns as you encounter them, rather than expecting to absorb everything at once. She encourages asking direct questions during the interview process itself, how the team codes, what's expected, requesting sample code, and pushes back on the instinct to stay quiet out of fear of looking inexperienced. Her manager's central challenge, she notes, is standardizing how a team with wildly different experience levels (from ten years in R to ten months) all code similarly enough that scripts remain genuinely shareable across the group.

Automating Recurring Reports: "Let Your Code Do Your Work for You"

Mansi's team has automated its recurring reporting workflow to the point where scripts for quarterly, weekly, and daily enrollment reports run unattended at 4am and email results directly to her and colleagues, requiring her only to forward the output to the appropriate regulatory office. Selecting and running the relevant script for a given patient cohort (adult, pregnancy, pediatric) produces around ten CSV reports with no manual coding required unless a quarter's protocol changes. Her guiding principle, drawing an explicit parallel to letting money make money, is to design code so it doesn't need to be rewritten every time a report is due.

Using AI as a Debugging Assistant, Not a Replacement

Mansi is direct that she does not use AI to write code for her, but relies on it heavily as a debugging and learning aid, pasting error messages into ChatGPT the same way she'd search Stack Overflow, a habit she traces back to a graduate-school professor's advice that a coder's core skill is knowing what to Google. She notes SAS's own error messages are often more explanatory and directly actionable than R's, which she finds comparatively unhelpful when something breaks. She draws an explicit analogy to Grammarly: a tool for catching and explaining mistakes, not one that writes the content for you, and warns that outsourcing actual coding to AI is a step too far ("just give them your salary in that case"). Recurring AI-flagged mistakes get logged into her personal cheat sheet so she stops needing to ask AI or search for the same fix twice. Separately, she and Tomás both note that despite SAS having comparatively little public training data available relative to Python, LLMs are still surprisingly accurate at producing SAS code.

CDISC Standards and Regulatory Compliance in Practice

Asked about regulatory expectations, Mansi notes compliance requirements vary by organization and data sensitivity, PHI and FDA-facing datasets carry more explicit requirements, and recommends early-career programmers proactively study CDISC standards (SDTM, ADaM) even if not deeply covered in graduate coursework; she supplemented her own education with certificate courses covering how to transform raw Excel or CSV data into CDISC-standard formats. Her broader advice is to keep updating that knowledge continuously, since the standards themselves keep evolving, and that the learning curve flattens considerably after hands-on practice (she cites having personally transformed 10 to 11 databases into standard formats as the point things "stopped feeling difficult").

Notable Quotes

"The code should tell a story of its own. The code should explain itself and it should be able to present what the author has put on the efforts on." — Mansi Thakrar
"It's very systematic in the way that if your code has to be run over and over again, make it as automated as possible... let your code do your work for you." — Mansi Thakrar
"I would say rather than using AI to do your job for you, use AI smartly where you can take it as assistance." — Mansi Thakrar
"If I keep making the same mistake again and again, I just take that and paste it in my cheat sheet... that helps me a lot to make sure that I'm not making the same mistakes again." — Mansi Thakrar
"Be ready for a lot of change. Be flexible. Try to learn as much as you can. No matter how much you do learn, your job is going to teach you 10 new things that you didn't know." — Mansi Thakrar

FAQ

What does Mansi consider good clinical programming practice?

Code that is functional, shareable, and self-documenting, meaning it explains its own purpose through structure (sections, indexing, comments) and reusable functions, so a colleague can understand, modify, or extend it without tracking down the original author.

How did Mansi's team automate its recurring clinical reports?

Scripts for quarterly, weekly, and daily enrollment reports run unattended overnight (around 4am), automatically producing CSV outputs and emailing them to the team, reducing what used to be manual reporting work to forwarding the output to the relevant regulatory office.

How does Mansi use AI tools like ChatGPT in her programming work?

Strictly as a debugging and learning assistant, pasting error messages to understand what went wrong, similar to searching Stack Overflow, rather than asking AI to generate code outright. She logs recurring fixes into a personal cheat sheet to reduce repeat lookups over time.

Why does Mansi say SAS is sometimes easier to debug than R?

She finds SAS's error messages more explanatory and actionable, often suggesting what the correct code should have been, whereas R's error messages, in her experience, explain the underlying issue less clearly.

What advice does Mansi give to programmers just starting out or joining a new team?

Expect the first three to six months to be disorienting as you learn a team's specific coding conventions; build a personal cheat sheet of recurring patterns and functions, ask direct questions without hesitation, and stay flexible, since every job teaches new things regardless of prior experience.

About The Verisian Community Podcast

The Verisian Community Podcast brings together experts in clinical trials to exchange innovative ideas and best practices central to clinical reporting, submission and review. Aligned with Verisian's mission to accelerate the evaluation and market launch of new medical treatments, each episode features expert insights, with guests ranging from statistical programmers to medical writers, to discuss the challenges and opportunities of the latest software and technology.

You can listen to us on YouTube, Apple Podcasts or Spotify.

Resources

‍

Explore more
Customer Stories
How Bayer Uses Verisian AI to Automate Submission Document Generation
Bayer used Verisian AI to automate key submission documentation for a non-CDISC clinical trial, producing submission-ready drafts in under two days and demonstrating how traceability-driven AI can dramatically reduce statistical programming busywork while maintaining regulatory quality.
Podcasts
The New Age Statistical Programmer
In this Verisian Community Podcast episode, we welcome Ritika Aggarwal from Novartis, who shares insights on the evolving landscape of statistical programming.
Tomás Sabat Stöfsel
March 14, 2024
Podcasts
The Limitations and Opportunities of Large Language Models
In this episode of the Verisian Community Podcast, we welcome Ricardo Baeza-Yates, a renowned computer scientist, to discuss the world of Large Language Models
Tomás Sabat Stöfsel
August 23, 2024