
Sponsors take months between the last patient last visit (LPLV) of a final pivotal trial and submission, time that sits directly on the critical path for time to market. These months are largely spent on the production and validation of analysis results and the generation of submission documents, before the final submission package is assembled.
This article sets out how sponsors prepare for submission today, why it takes as long as it does, and how leaders can set up their organisations to cut months from time to market by enabling submission within days after LPLV. The pressure to accelerate timelines has acquired recent urgency. Earlier this year, the FDA announced two proof-of-concept trials reporting endpoints and safety signals to the agency in real time. The time is right to rethink how sponsors become submission-ready. The EMA is also running a pilot to determine how to manage raw clinical data.
By the time a programme reaches its final pivotal trial, most of the science is already done. The hypotheses and endpoints were set in earlier phases. The analysis plan is signed. The dose was settled in Phases 1 and 2. What the last trial supplies is confirmation. The efficacy summary takes its shape from the earlier studies; the pivotal trial confirms it or does not. The safety picture is much the same: the risks are identified early and largely confirmed at the end, with late-breaking and rare events as the additions.
Much of what goes into a submission is therefore known, and could be assembled, before the final pivotal trial reads out. Some key outputs nonetheless depend on the final data: analyses and results must be run and validated, and the final submission documents produced from them.
Sponsors have invested substantially in accelerating this process, and timelines are still not what they should be. The reason is that the penalty for error makes caution rational: a submission that arrives with too many errors comes back on GCP and data quality grounds. Organisations therefore trust nothing until everything is final. No output can be acted on before lock, and clinical and medical leaders will not engage until the numbers can be trusted.
The same instinct shapes the analysis plan. Regulators treat a signed SAP as a commitment, and an adjustment made after seeing the data invites the suspicion that it was made because of the data, so the time to argue about the analysis is before the SAP is signed off. Teams tend to do the reverse, rushing the plan and relitigating it once the numbers land, which settles the analysis late and everything downstream later still.
The same caution governs the handoffs between departments. Data management collects and cleans the data, biometrics derives the analyses and outputs from it, medical writing turns those into the clinical study report and summaries, and regulatory affairs assembles the package. Statistical programmers routinely find data quality issues that require data management to make changes, sometimes after lock, and because reopening a locked database invites its own problems, the correction is often hard-coded in the analysis instead. That is how the data, the code and the documents begin to diverge. Compressing one department's timeline without changing this moves the risk along rather than removing it: if data management is squeezed, the pressure lands on regulatory operations, where a gateway failure means no submission at all.
Underneath all of this is the volume of material that has to be consistent. Novo Nordisk's former Chief Scientific Officer once observed that the paperwork for two insulin products, printed and stacked, would stand taller than the Empire State Building. The problem is not the volume itself but that all of its contents have to be consistent: the specifications with the analysis plan, the code with the specifications, the outputs with the code, and the Define-XML, Reviewer's Guides and summaries with the outputs they describe. Regulatory affairs can confirm a document is complete, but it cannot confirm that a figure is the one the analysis produced, because the evidence sits in code that regulatory affairs does not read. Keeping all of that consistent as a study changes is manual work with no process of its own, it appears in no plan, estimate or budget line, and it is the work that consumes the months.
The industry has already worked through some of these problems in one function. In the CMC space, regulators moved away from flat documents toward data they can parse and interrogate, and the quality function made that shift under quality by design: the process is designed to produce a quality product rather than having quality confirmed by testing finished batches. Clinical has not made the equivalent change to its underlying business process.
Part of the reason is that the submission process defined by regulators and industry over the last 40 years is document-focused: signature and traceability still work largely in document terms. Its units are metadata, supporting documents, labelling and registration, and its measure of readiness is whether the documents are complete. Regulators beyond the FDA increasingly want the data itself, meaning the datasets, the derivations and the ability to interrogate what was done.
Every reason above comes back to the same point: nobody can act on an output until it has been shown to be right, and what produces every output is code. The protocol defines the study design, the SAP defines the analysis, the specifications define required variables and datasets, and all of it is implemented in the analysis programs that turn the raw collected data into the results. Those programs are the only artefact that actually produced the numbers; the documents describe what the programs were meant to do, or what someone believes they did. If the code is right, the results are right, and a document is right only to the extent that it matches the code. That is why quality is the lever rather than speed, and why validation of the code is where the answer starts.
In practice the path from LPLV to submission runs as follows. Reaching database lock itself takes weeks: draft SDTM and ADaM datasets and draft outputs are produced and checked, the study is unblinded, and if nothing is found, which rarely happens, the database is locked. The final readouts follow, and the integrated safety and efficacy databases are built from them. Only when all of that is in order does the work that sits most directly on the critical path begin: the Define-XML for SDTM and for ADaM, the SDRG and the ADRG, and the documents that describe how each result was produced.
Validation is what allows each step to proceed. Validated analyses are analyses that can be trusted; trusted analyses are what the submission documents describe; and documents generated from validated analyses are what give regulatory affairs, and the reviewer, confidence that a figure is the one the analysis produced. Ultimately, it is the approval that counts, not the submission date, and a filing that arrives earlier and generates enough questions to consume the gain has achieved nothing. If validation runs continuously as the study proceeds, each of those steps is already done when LPLV arrives, and the sequence that consumes the months no longer has to run in series.
The solution is a study that is submission-ready at all times. That means analyses and derivations validated continuously against their specifications as the study runs, so that a change to a specification or to the code is checked when it happens rather than at lock. It means every reported number is traceable back to the source data that produced it, so that a reviewer inside the sponsor or at the agency can establish how a figure was derived without reconstructing it. And it means submission documentation generated from that same validated code rather than written separately, so that the document and the analysis cannot disagree.
The system that does this has to be deterministic at every point where evidence must hold. Deterministic means the same inputs produce the same output every time, by fixed rules, and the result can be reproduced and audited. Probabilistic means a model infers meaning, including from prose such as a SAP, and may produce different output on different runs; that is useful for proposing a discrepancy for a person to judge, and not acceptable as the final arbiter of whether an analysis matches what was pre-specified. A naive AI system that generates code and submission documents from a prompt is unacceptable on quality grounds and will not be accepted by regulators or by QA. AI applied over this process without a human in the loop will not be accepted, and should not be.
Shortened timelines are the consequence, not the objective. A study that is validated continuously does not need the months after LPLV, because the work those months contained has already been done, and the filing follows from the readout rather than beginning with it. The programming workload is what currently makes the sequence serial, and removing it from the critical path is what brings submission to within days of LPLV.
Moira Daniels is a regulatory strategist with twenty-five years in regulatory affairs, combining commercial and technical regulatory experience in complex regulatory environments. She was most recently Senior Vice President and Head of Regulatory EMEA, International and Operations at BridgeBio.
Andrew York has worked in statistical programming for nearly forty years, more than thirty of them in management, most recently as Vice President of Clinical Data Science at Novo Nordisk. Over that period he has contributed to the adoption of double programming, the implementation of CDISC standards, and the transitions to R and to the use of artificial intelligence in the discipline. He is a strategic advisor to Verisian.
Tomás Sabat Stöfsel is the CEO and co-founder of Verisian. He was previously COO and a founding team member of TypeDB, the open-source database, where he spent six years building open-source software for drug discovery and development. He is a graduate of the University of Cambridge and has spent the past decade founding and building technology businesses.
