.png)
Andrew York has over 30 years experience as a statistical programmer and biostatistician. From the beginning of his career as a statistician at Roche, to his current role as Clinical Data Science Vice President at Novo Nordisk, Andy has seen the industry evolve and has played a leading role in shaping best practice in the industry.
Reflecting on these experiences, in this episode Andy shares his unique perspectives on the evolution of statistical programmers. Alongside this, Tomás and Andy discuss open-source programming, where the industry would be without standards, and the need for standards that are more sympathetic to AI and automation driven processes. Andy also explains why achieving full end-to-end traceability currently poses such a challenge, and why doing so has immense benefits to changing how we program studies today.
In this conversation, we explore:
Guest: Andrew York, Clinical Data Scientist Vice President, Novo Nordisk, a statistical programmer for four decades, spanning SmithKline & French, Roche, Covance, and Novo Nordisk.
Host: Tomás Sabat Stöfsel, CEO and Co-Founder, Verisian
Topics: History of statistical programming, double programming origins, CDISC/SDTM/ADaM standards, open source (R, Python, SAS), data traceability, AI in clinical programming, the future of the clinical data scientist role
The full episode can be found on YouTube, Apple Podcasts or Spotify.
Double programming, having two independent programmers verify the same output, is not a regulatory mandate. It emerged in the 1990s as a faster alternative to the fully manual traceability process used until then, where a second programmer manually reviewed every step of a colleague's code and output line by line. Andrew York, who witnessed its introduction firsthand, explains that once one CRO adopted the practice and saw efficiency gains, competitive pressure spread it across the industry, first among CROs, then into pharma companies directly. It was, in his words, a way of "reducing the amount of effort that you took to get a reasonable assurance of quality," not a guarantee against error, since two programmers can still independently make the same mistake.
CDISC's SDTM standard was widely and consistently adopted because its structure, one variable per column, patient and visit data running down the rows, closely matches how clinical data naturally sits in a database. ADaM, developed in the early 2000s, took a fundamentally different, "rotated" structure designed to make analysis "one PROC away" from a finished table. Andrew York says that promise has never been fully realized in practice: clinical data complexity means extra derivation steps are usually still required, and ADaM remains formally a guideline rather than a strict standard, with implementation varying by company. He attributes this not to ADaM being a failure, but to an inherent tension between standardization and the exploratory nature of clinical science.
Novo Nordisk takes a "right language for the right task" approach rather than mandating a single statistical computing environment, and has submitted its own R packages to the public domain at the FDA's specific request. Andrew York argues R is not truly "free" once cloud infrastructure costs are counted, but acknowledges SAS now faces real competitive pressure it didn't have a decade ago. He believes SAS will remain necessary for at least the next decade because of the volume of legacy code and data built up over 40 years of industry use, even as R adoption accelerates, driven partly by a generational shift (his estimate: the average SAS programmer is in their late 30s, the average R programmer in their 20s). Python remains further behind for regulatory submissions specifically because, in his account, "the regulators are saying, please don't", despite Novo Nordisk having already run at least one fully validated safety analysis entirely in Python.
For Andrew York, traceability means being able to demonstrate exactly how raw clinical trial data became a final table, and it's the real mechanism that could eventually reduce or replace double programming. He argues CDISC's traceability variables provide only a "checkbox" level of assurance compared to genuine, automated, code-level traceability. Notably, he states that metadata "doesn't tell you whether a drug works or not", the underlying data and code are what matter, with metadata serving only as a layer of documentation on top. This distinction, traceability at the code level versus traceability at the standards/metadata level, is a recurring theme in the conversation and directly informed why Verisian was founded.
Andrew York is direct about where current AI falls short: a Virginia Tech study found one-shot LLM-generated statistical analysis code was accurate only about 51% of the time on comparatively simple analyses, well below the industry's quality expectations. He compares this to output he's seen from AI code-generation vendors, describing the required text-to-code process as too slow and too simplistic for programs that can run over 1,000 lines. His near-term prediction is agentic AI systems that handle traceability and flag discrepancies for a human programmer to review, shifting the statistical programmer role toward what Novo Nordisk already calls a "clinical data scientist," focused on interpreting data and the story it tells rather than hand-writing code. He expects this shift to unfold over the next 2-10 years, with hands-on programming skills remaining relevant for "a number of years" yet.
"Nobody is telling us that we have to do double programming. What they're telling us is that we have to create something with quality." — Andrew York
"If you start out with a tool that does not have quality from day one, it's not starting at 100% quality. It's starting more like at 50% or lower, in my view." — Andrew York
"The actual result is the most important thing. It doesn't matter, the metadata doesn't tell you whether a drug works or not." — Andrew York
"We definitely want to move into a world where the AI is helping the programmer understand what's going on, but doing a lot of the heavy lifting on their behalf." — Andrew York
"In the future, a statistical programmer may well become, and as we call them in Novo, clinical data scientists, the emphasis is on understanding the data, not so much understanding how we got there with the code." — Andrew York
Why does double programming exist in clinical statistical programming?
It emerged in the 1990s as a faster, less resource-intensive replacement for fully manual traceability checks, then spread across the industry through competitive adoption, not because it was ever formally mandated by regulators.
What's the difference between SDTM and ADaM standards?
SDTM organizes data with one variable per column, closely matching natural database structure; this drove consistent industry-wide adoption. ADaM restructures data to make analysis theoretically "one PROC away" from a final table, but is a guideline rather than a strict standard, and clinical data complexity means that promise isn't always realized in practice.
How accurate are large language models at writing clinical statistical analysis code?
A Virginia Tech study cited in the episode found one-shot LLM outputs for simple statistical analyses were accurate only around 51% of the time, a key reason the guest argues LLMs cannot yet safely replace human programmers in regulated pharma work.
Will AI replace statistical programmers in the pharmaceutical industry?
Not according to Andrew York, he expects AI's near-term value to come from automating traceability and flagging discrepancies for human review, not from generating submission-ready code autonomously, with the programmer role evolving into "clinical data scientist" rather than disappearing.
Why is Novo Nordisk submitting R packages to the public domain?
Because the FDA specifically told Novo it would prefer their packages be publicly available rather than kept internal, reflecting regulators' broader preference for community-validated, open-source statistical tools over proprietary code.
The Verisian Community Podcast brings together experts in clinical trials to exchange innovative ideas and best practices central to clinical reporting, submission and review. Aligned with Verisian's mission to accelerate the evaluation and market launch of new medical treatments, each episode features expert insights, with guests ranging from statistical programmers to medical writers, to discuss the challenges and opportunities of the latest software and technology.
You can listen to us on YouTube, Apple Podcasts or Spotify.