We use cookies to ensure that we give you the best experience on our website.
‍Read about cookies preferences.
by
Tomás Sabat Stöfsel
September 11, 2025
1 min
The Evolution of Statistical Programming: How AI, Open Source & Traceability are Shaping the Future

Episode Summary

Andrew York has over 30 years experience as a statistical programmer and biostatistician. From the beginning of his career as a statistician at Roche, to his current role as Clinical Data Science Vice President at Novo Nordisk, Andy has seen the industry evolve and has played a leading role in shaping best practice in the industry.

Reflecting on these experiences, in this episode Andy shares his unique perspectives on the evolution of statistical programmers. Alongside this, Tomás and Andy discuss open-source programming, where the industry would be without standards, and the need for standards that are more sympathetic to AI and automation driven processes. Andy also explains why achieving full end-to-end traceability currently poses such a challenge, and why doing so has immense benefits to changing how we program studies today.

In this conversation, we explore:

  • The history of statistical programming and how the modern statistical programmer came to be
  • Standards and their role in driving the industry forward
  • Opportunities and challenges of open source programming
  • Traceability and its potential in changing how we build and validate programs today
  • The future of statistical programming and how AI will transform current ways of working

Key Takeaways

  • Double programming exists because manual traceability in 1990s CROs was too resource-intensive to verify data accuracy any other way, it was never a regulatory requirement, just an emergent industry best practice.
  • CDISC's SDTM standard succeeded because it mirrors how data is naturally structured (one variable per column); ADaM's "one PROC away" promise has been only partially realized because clinical data complexity resists full standardization.
  • Novo Nordisk's own submission-grade R packages were made public at the FDA's request, reflecting a broader industry shift toward open-source statistical computing (R and, increasingly, Python) alongside SAS rather than replacing it outright.
  • Current one-shot LLM accuracy on statistical analysis code is around 51%, according to a Virginia Tech study Andrew and host Tomás discuss, a key reason large language models cannot yet safely replace human programmers in a regulated environment.
  • Andrew York predicts AI's near-term value lies in automated traceability and QC flagging, not autonomous code generation, with statistical programmers eventually evolving into "clinical data scientists" who interpret data rather than hand-write code.

Episode Info

Guest: Andrew York, Clinical Data Scientist Vice President, Novo Nordisk, a statistical programmer for four decades, spanning SmithKline & French, Roche, Covance, and Novo Nordisk.

Host: Tomás Sabat Stöfsel, CEO and Co-Founder, Verisian

Topics: History of statistical programming, double programming origins, CDISC/SDTM/ADaM standards, open source (R, Python, SAS), data traceability, AI in clinical programming, the future of the clinical data scientist role

The full episode can be found on YouTube, Apple Podcasts or Spotify.

How Double Programming Actually Started

Double programming, having two independent programmers verify the same output, is not a regulatory mandate. It emerged in the 1990s as a faster alternative to the fully manual traceability process used until then, where a second programmer manually reviewed every step of a colleague's code and output line by line. Andrew York, who witnessed its introduction firsthand, explains that once one CRO adopted the practice and saw efficiency gains, competitive pressure spread it across the industry, first among CROs, then into pharma companies directly. It was, in his words, a way of "reducing the amount of effort that you took to get a reasonable assurance of quality," not a guarantee against error, since two programmers can still independently make the same mistake.

The Promise and Reality of CDISC Standards (SDTM and ADaM)

CDISC's SDTM standard was widely and consistently adopted because its structure, one variable per column, patient and visit data running down the rows, closely matches how clinical data naturally sits in a database. ADaM, developed in the early 2000s, took a fundamentally different, "rotated" structure designed to make analysis "one PROC away" from a finished table. Andrew York says that promise has never been fully realized in practice: clinical data complexity means extra derivation steps are usually still required, and ADaM remains formally a guideline rather than a strict standard, with implementation varying by company. He attributes this not to ADaM being a failure, but to an inherent tension between standardization and the exploratory nature of clinical science.

Open Source, R, Python, and the Future of SAS

Novo Nordisk takes a "right language for the right task" approach rather than mandating a single statistical computing environment, and has submitted its own R packages to the public domain at the FDA's specific request. Andrew York argues R is not truly "free" once cloud infrastructure costs are counted, but acknowledges SAS now faces real competitive pressure it didn't have a decade ago. He believes SAS will remain necessary for at least the next decade because of the volume of legacy code and data built up over 40 years of industry use, even as R adoption accelerates, driven partly by a generational shift (his estimate: the average SAS programmer is in their late 30s, the average R programmer in their 20s). Python remains further behind for regulatory submissions specifically because, in his account, "the regulators are saying, please don't", despite Novo Nordisk having already run at least one fully validated safety analysis entirely in Python.

Why Traceability Matters

For Andrew York, traceability means being able to demonstrate exactly how raw clinical trial data became a final table, and it's the real mechanism that could eventually reduce or replace double programming. He argues CDISC's traceability variables provide only a "checkbox" level of assurance compared to genuine, automated, code-level traceability. Notably, he states that metadata "doesn't tell you whether a drug works or not", the underlying data and code are what matter, with metadata serving only as a layer of documentation on top. This distinction, traceability at the code level versus traceability at the standards/metadata level, is a recurring theme in the conversation and directly informed why Verisian was founded.

AI's Realistic Role in the Future of Clinical Data Processing

Andrew York is direct about where current AI falls short: a Virginia Tech study found one-shot LLM-generated statistical analysis code was accurate only about 51% of the time on comparatively simple analyses, well below the industry's quality expectations. He compares this to output he's seen from AI code-generation vendors, describing the required text-to-code process as too slow and too simplistic for programs that can run over 1,000 lines. His near-term prediction is agentic AI systems that handle traceability and flag discrepancies for a human programmer to review, shifting the statistical programmer role toward what Novo Nordisk already calls a "clinical data scientist," focused on interpreting data and the story it tells rather than hand-writing code. He expects this shift to unfold over the next 2-10 years, with hands-on programming skills remaining relevant for "a number of years" yet.

Notable Quotes

"Nobody is telling us that we have to do double programming. What they're telling us is that we have to create something with quality." — Andrew York
"If you start out with a tool that does not have quality from day one, it's not starting at 100% quality. It's starting more like at 50% or lower, in my view." — Andrew York
"The actual result is the most important thing. It doesn't matter, the metadata doesn't tell you whether a drug works or not." — Andrew York
"We definitely want to move into a world where the AI is helping the programmer understand what's going on, but doing a lot of the heavy lifting on their behalf." — Andrew York
"In the future, a statistical programmer may well become, and as we call them in Novo, clinical data scientists, the emphasis is on understanding the data, not so much understanding how we got there with the code." — Andrew York

FAQ

Why does double programming exist in clinical statistical programming?

It emerged in the 1990s as a faster, less resource-intensive replacement for fully manual traceability checks, then spread across the industry through competitive adoption, not because it was ever formally mandated by regulators.

What's the difference between SDTM and ADaM standards?

SDTM organizes data with one variable per column, closely matching natural database structure; this drove consistent industry-wide adoption. ADaM restructures data to make analysis theoretically "one PROC away" from a final table, but is a guideline rather than a strict standard, and clinical data complexity means that promise isn't always realized in practice.

How accurate are large language models at writing clinical statistical analysis code?

A Virginia Tech study cited in the episode found one-shot LLM outputs for simple statistical analyses were accurate only around 51% of the time, a key reason the guest argues LLMs cannot yet safely replace human programmers in regulated pharma work.

Will AI replace statistical programmers in the pharmaceutical industry?

Not according to Andrew York, he expects AI's near-term value to come from automating traceability and flagging discrepancies for human review, not from generating submission-ready code autonomously, with the programmer role evolving into "clinical data scientist" rather than disappearing.

Why is Novo Nordisk submitting R packages to the public domain?

Because the FDA specifically told Novo it would prefer their packages be publicly available rather than kept internal, reflecting regulators' broader preference for community-validated, open-source statistical tools over proprietary code.

About The Verisian Community Podcast

The Verisian Community Podcast brings together experts in clinical trials to exchange innovative ideas and best practices central to clinical reporting, submission and review. Aligned with Verisian's mission to accelerate the evaluation and market launch of new medical treatments, each episode features expert insights, with guests ranging from statistical programmers to medical writers, to discuss the challenges and opportunities of the latest software and technology.

You can listen to us on YouTube, Apple Podcasts or Spotify.

Resources

‍

Explore more