
In this episode, Shafi discusses his decades-long journey and what led him into the world of statistical programming. He shares insights into the evolution of automation in clinical trials, its role in transforming traditional tasks like validation and double programming, and the impact AI is having on oversight and traceability for large teams. Shafi also addresses the ongoing challenges in scaling AI for clinical trial reporting and how embracing these technologies is critical for staying ahead in the evolving landscape of the industry.
Guest: Shafi Chowdhury, founder of Shafi Consultancy, statistical programmer with close to three decades of industry experience.
Host: Tomás Sabat Stöfsel, Verisian Community Podcast
Topics: How Shafi Consultancy began from informally training friends and relatives, building trust through transparency as a consulting philosophy, why TLF production has barely changed in 30 years, metadata-driven e-templates and 80% TLF auto-generation, good versus bad automation design, AI's near-term role in table of contents and ADaM generation, the evolving role of the statistical programmer, remote work
The full episode can be found on YouTube, Apple Podcasts or Spotify.
Shafi Chowdhury describes Shafi Consultancy's origin as accidental. It began with him informally teaching friends and relatives to program in SAS after work, then introducing his strongest students to clients who were struggling to find programmers, initially at no cost, until the clients saw the results and started hiring them properly. The same model scaled when he started a second office in Bangladesh, spending the first year almost entirely on training before the team became self-sufficient. Now 23 years in, Shafi credits the consultancy's growth less to technical differentiation and more to a deliberate philosophy of sharing everything he learns rather than guarding it, arguing that transparency is what builds the trust a consulting relationship actually depends on, and that visibly withholding knowledge, even the appearance of it, can quietly damage a client relationship faster than almost anything else.
Shafi's most persistent frustration is that the manual, output-by-output process for producing safety and efficacy tables, listings, and figures has barely changed since his 1993 placement year, despite three decades of CDISC standardization. His view is that CDISC "is just another standard," useful, but not by itself a driver of automation, since the underlying manual production process it standardizes has stayed largely the same. He's more optimistic about CDISC's newer results datasets, which he sees as finally opening the door to the kind of automation his team has been building toward for years.
Around eight years ago, Shafi's team built a results-database-driven system using what he calls "e-templates," essentially structured, Excel-style specifications where a user selects the analysis they want and the underlying program auto-generates, reaching a first working version roughly six years ago. His reasoning is that most of the actual time spent on TLF programming goes into making output look presentable, not into producing the underlying statistical result itself, which is often just a single procedure call. With that layer automated, his team can now auto-generate around 80% of TLFs directly from metadata, producing consistent, identically structured output regardless of which programmer or location is involved.
Shafi draws a sharp distinction between automation done well and done badly. Good automation is modular and defaults-driven, so the end user, in his coffee-and-pizza analogy, gets "coffee with your milk," not a menu of ten milk options to weigh before every table. Bad automation looks like giant, opaque macros: he recounts a client whose macro required setting over 100 parameters per table, where a single wrong selection could silently produce incorrect results that got waved through anyway because "it's from a validated macro." His prescription is that automation should be judged by whether it saves time and builds trust in the numbers it produces, not by how technically impressive the underlying code is to fellow programmers.
Shafi is skeptical of individual programmers using ChatGPT to generate one-off code, arguing the time saved is often eaten up by having to reverse-engineer unfamiliar output, "it produced output... is it correct? that's often the tricky part." His higher-value near-term use case is table of contents and table shell generation: pharma companies already have years of prior TOCs and shells for similar trial designs, making this a well-scoped training problem. He expects AI-generated TOCs to start around 50% complete and improve to 90 to 95% with statistician review, with table shells following a similar trajectory, and projects AI reaching roughly 80% automation of general programming work over time. He connects this directly to freeing statistical programmers from repetitive dataset and demography table work so they can spend time on genuinely complex endpoints and validation work that requires deep attention, the kind of work pharma companies currently decline simply because programmers don't have the time.
"It still frustrates me that what I did in my placement year, we're talking like 1993 here, right? This is 31 years ago. What I did then, we are still producing now." — Shafi Chowdhury
"Automation should always be about saving time. It should be about efficiency, it should be about quality. That actually they can trust their numbers that are coming out." — Shafi Chowdhury
"It's not about your skill, it's how easy can you make it for the user, right?" — Shafi Chowdhury
"We're producing PDF reports, which is nothing more than a glorified piece of paper that we used to produce 30 years ago. So we need to think bigger." — Shafi Chowdhury
"The truth is always somewhere in the middle... you have to find for your own company what the individual challenges are." — Shafi Chowdhury
Why does Shafi Chowdhury say TLF production hasn't really changed in 30 years?
Because the manual, table-by-table process he used in his 1993 placement year is still the industry default today. In his view, CDISC standardized data formats but didn't itself drive automation of the underlying production process, though newer CDISC results datasets are starting to change that.
What is a metadata-driven e-template system for TLF generation?
An approach Shafi's team built roughly eight years ago where a user specifies the desired analysis through an Excel-style template, and the analysis program is auto-generated from a results database rather than hand-coded, reaching around 80% TLF auto-generation for standard outputs.
What's the difference between good and bad automation, according to Shafi?
Good automation is modular and default-driven, requiring minimal input from ordinary users. Bad automation is an opaque macro with dozens or hundreds of manually set parameters, where a single wrong choice can silently produce incorrect results that get accepted because they came from a "validated" tool.
What can AI realistically automate in clinical statistical programming today?
Shafi's highest-confidence near-term use case is table of contents and table shell generation from trial design inputs, projecting roughly 90 to 95% completion with statistician review, versus around 80% automation of general programming work over a longer time horizon. He's skeptical of ad hoc ChatGPT-generated code for individual analyses.
How does Shafi Chowdhury think the statistical programmer's role should evolve?
Away from repetitive, duplicated dataset and demography table work and toward genuinely complex endpoint programming and validation, work that pharma companies currently decline to take on simply because their programmers lack the time, not the skill.
The Verisian Community Podcast brings together experts in clinical trials to exchange innovative ideas and best practices central to clinical reporting, submission and review. Aligned with Verisian's mission to accelerate the evaluation and market launch of new medical treatments, each episode features expert insights, with guests ranging from statistical programmers to medical writers, to discuss the challenges and opportunities of the latest software and technology.
You can listen to us on YouTube, Apple Podcasts or Spotify.
