We use cookies to ensure that we give you the best experience on our website.
‍Read about cookies preferences.
by
Tomás Sabat Stöfsel
October 8, 2024
1 min read
Automation and AI in Clinical Trial Reporting

Episode Summary

In this episode, Shafi discusses his decades-long journey and what led him into the world of statistical programming. He shares insights into the evolution of automation in clinical trials, its role in transforming traditional tasks like validation and double programming, and the impact AI is having on oversight and traceability for large teams. Shafi also addresses the ongoing challenges in scaling AI for clinical trial reporting and how embracing these technologies is critical for staying ahead in the evolving landscape of the industry.

Key Takeaways

  • Shafi Chowdhury still finds it "depressing" that the safety and basic efficacy TLF programs he wrote during his 1993 placement year are produced in essentially the same manual way today, 31 years and multiple CDISC standard revisions later.
  • Shafi Consultancy started by accident: Shafi began informally teaching friends and relatives to program in SAS after work, then began placing them with his own clients, a model built entirely on sharing knowledge rather than protecting it as a competitive advantage.
  • Roughly eight years ago, Shafi's team built a metadata-driven, e-template-based system that can auto-generate about 80% of standard TLF programs, arguing most manual programming time is spent making output "look nice" rather than producing the actual analysis result.
  • Good automation, in Shafi's framing, is modular and defaults-driven so ordinary users never have to touch it; bad automation is a giant, black-box macro with 100+ parameters where a single wrong selection silently produces incorrect numbers that get accepted anyway because "it's from a validated macro."
  • Shafi expects AI to reach roughly 95% automation of table of contents and table shell creation and around 80% of programming work over time, framing the near-term highest-value use case as generating TOCs and shells from trial design inputs rather than having individual programmers prompt ChatGPT for one-off code.

Episode Info

Guest: Shafi Chowdhury, founder of Shafi Consultancy, statistical programmer with close to three decades of industry experience.

Host: Tomás Sabat Stöfsel, Verisian Community Podcast

Topics: How Shafi Consultancy began from informally training friends and relatives, building trust through transparency as a consulting philosophy, why TLF production has barely changed in 30 years, metadata-driven e-templates and 80% TLF auto-generation, good versus bad automation design, AI's near-term role in table of contents and ADaM generation, the evolving role of the statistical programmer, remote work

The full episode can be found on YouTube, Apple Podcasts or Spotify.

From Teaching Friends to a 23-Year Consultancy

Shafi Chowdhury describes Shafi Consultancy's origin as accidental. It began with him informally teaching friends and relatives to program in SAS after work, then introducing his strongest students to clients who were struggling to find programmers, initially at no cost, until the clients saw the results and started hiring them properly. The same model scaled when he started a second office in Bangladesh, spending the first year almost entirely on training before the team became self-sufficient. Now 23 years in, Shafi credits the consultancy's growth less to technical differentiation and more to a deliberate philosophy of sharing everything he learns rather than guarding it, arguing that transparency is what builds the trust a consulting relationship actually depends on, and that visibly withholding knowledge, even the appearance of it, can quietly damage a client relationship faster than almost anything else.

Why We're Still Producing the Same TLFs as 1993

Shafi's most persistent frustration is that the manual, output-by-output process for producing safety and efficacy tables, listings, and figures has barely changed since his 1993 placement year, despite three decades of CDISC standardization. His view is that CDISC "is just another standard," useful, but not by itself a driver of automation, since the underlying manual production process it standardizes has stayed largely the same. He's more optimistic about CDISC's newer results datasets, which he sees as finally opening the door to the kind of automation his team has been building toward for years.

Metadata-Driven E-Templates: Automating 80% of TLF Production

Around eight years ago, Shafi's team built a results-database-driven system using what he calls "e-templates," essentially structured, Excel-style specifications where a user selects the analysis they want and the underlying program auto-generates, reaching a first working version roughly six years ago. His reasoning is that most of the actual time spent on TLF programming goes into making output look presentable, not into producing the underlying statistical result itself, which is often just a single procedure call. With that layer automated, his team can now auto-generate around 80% of TLFs directly from metadata, producing consistent, identically structured output regardless of which programmer or location is involved.

Good Automation Is Invisible, Bad Automation Is a 100-Parameter Black Box

Shafi draws a sharp distinction between automation done well and done badly. Good automation is modular and defaults-driven, so the end user, in his coffee-and-pizza analogy, gets "coffee with your milk," not a menu of ten milk options to weigh before every table. Bad automation looks like giant, opaque macros: he recounts a client whose macro required setting over 100 parameters per table, where a single wrong selection could silently produce incorrect results that got waved through anyway because "it's from a validated macro." His prescription is that automation should be judged by whether it saves time and builds trust in the numbers it produces, not by how technically impressive the underlying code is to fellow programmers.

Where AI Actually Helps Right Now: TOCs and Table Shells, Not One-Off Code

Shafi is skeptical of individual programmers using ChatGPT to generate one-off code, arguing the time saved is often eaten up by having to reverse-engineer unfamiliar output, "it produced output... is it correct? that's often the tricky part." His higher-value near-term use case is table of contents and table shell generation: pharma companies already have years of prior TOCs and shells for similar trial designs, making this a well-scoped training problem. He expects AI-generated TOCs to start around 50% complete and improve to 90 to 95% with statistician review, with table shells following a similar trajectory, and projects AI reaching roughly 80% automation of general programming work over time. He connects this directly to freeing statistical programmers from repetitive dataset and demography table work so they can spend time on genuinely complex endpoints and validation work that requires deep attention, the kind of work pharma companies currently decline simply because programmers don't have the time.

Notable Quotes

"It still frustrates me that what I did in my placement year, we're talking like 1993 here, right? This is 31 years ago. What I did then, we are still producing now." — Shafi Chowdhury
"Automation should always be about saving time. It should be about efficiency, it should be about quality. That actually they can trust their numbers that are coming out." — Shafi Chowdhury
"It's not about your skill, it's how easy can you make it for the user, right?" — Shafi Chowdhury
"We're producing PDF reports, which is nothing more than a glorified piece of paper that we used to produce 30 years ago. So we need to think bigger." — Shafi Chowdhury
"The truth is always somewhere in the middle... you have to find for your own company what the individual challenges are." — Shafi Chowdhury

FAQ

Why does Shafi Chowdhury say TLF production hasn't really changed in 30 years?

Because the manual, table-by-table process he used in his 1993 placement year is still the industry default today. In his view, CDISC standardized data formats but didn't itself drive automation of the underlying production process, though newer CDISC results datasets are starting to change that.

What is a metadata-driven e-template system for TLF generation?

An approach Shafi's team built roughly eight years ago where a user specifies the desired analysis through an Excel-style template, and the analysis program is auto-generated from a results database rather than hand-coded, reaching around 80% TLF auto-generation for standard outputs.

What's the difference between good and bad automation, according to Shafi?

Good automation is modular and default-driven, requiring minimal input from ordinary users. Bad automation is an opaque macro with dozens or hundreds of manually set parameters, where a single wrong choice can silently produce incorrect results that get accepted because they came from a "validated" tool.

What can AI realistically automate in clinical statistical programming today?

Shafi's highest-confidence near-term use case is table of contents and table shell generation from trial design inputs, projecting roughly 90 to 95% completion with statistician review, versus around 80% automation of general programming work over a longer time horizon. He's skeptical of ad hoc ChatGPT-generated code for individual analyses.

How does Shafi Chowdhury think the statistical programmer's role should evolve?

Away from repetitive, duplicated dataset and demography table work and toward genuinely complex endpoint programming and validation, work that pharma companies currently decline to take on simply because their programmers lack the time, not the skill.

About The Verisian Community Podcast

The Verisian Community Podcast brings together experts in clinical trials to exchange innovative ideas and best practices central to clinical reporting, submission and review. Aligned with Verisian's mission to accelerate the evaluation and market launch of new medical treatments, each episode features expert insights, with guests ranging from statistical programmers to medical writers, to discuss the challenges and opportunities of the latest software and technology.

You can listen to us on YouTube, Apple Podcasts or Spotify.

Resources

‍

Explore more
Customer Stories
How Bayer Uses Verisian AI to Automate Submission Document Generation
Bayer used Verisian AI to automate key submission documentation for a non-CDISC clinical trial, producing submission-ready drafts in under two days and demonstrating how traceability-driven AI can dramatically reduce statistical programming busywork while maintaining regulatory quality.
Podcasts
The New Age Statistical Programmer
In this Verisian Community Podcast episode, we welcome Ritika Aggarwal from Novartis, who shares insights on the evolving landscape of statistical programming.
Tomás Sabat Stöfsel
March 14, 2024
Podcasts
The Limitations and Opportunities of Large Language Models
In this episode of the Verisian Community Podcast, we welcome Ricardo Baeza-Yates, a renowned computer scientist, to discuss the world of Large Language Models
Tomás Sabat Stöfsel
August 23, 2024