The statistician sees the protocol for the first time six months after first patient in.
The primary endpoint reads “change from baseline in the symptom score.” There’s no visit window. The eCRF records whether rescue medication was used but not when, so there’s no clean way to handle patients who took it before the primary assessment. And nobody wrote down what “baseline” means for patients who were screened twice.
Each of these is fixable. Each fix now costs a protocol amendment, an eCRF change, a retraining cycle and months of data collected the old way.
If you only bring in biostatistics consulting services for the analysis, you’re hiring someone to interpret decisions they had no part in making. The work that decides whether your trial can answer its question happens before enrollment, and so should the involvement of your biostatistics partner. This guide covers what that partner should own before the first patient is dosed, and what to ask before you sign.
Why Biostatistics Belongs Before First Patient In
None of that can happen if the statistician arrives after the protocol is signed.
So what does a biostatistician do before a clinical trial? They start with the question the trial is trying to answer. The data come later. Endpoints, comparators, visit schedules and what counts as a missing value are all statistical decisions, whether or not a statistician makes them. That’s the real role of biostatistics in clinical trial design, and it’s why biostatistics for clinical trials is a planning job before it’s an analysis job.
A biostatistics partner should own, or co-own, a defined set of deliverables before first patient in. Here’s what that list usually looks like, and what goes wrong when each item arrives late.Pre-study deliverable What the partner owns What goes wrong if it’s late Endpoint and estimand definition Precise definition of what is measured, when, and how intercurrent events are handled Primary analysis can’t be computed as intended; mid-study amendments Sample size and power Assumptions, justification, sensitivity to those assumptions Underpowered trial, or an oversized and overpriced one Randomization and stratification Scheme, block sizes, stratification factors, specifications for the IxRS Imbalance, or strata that can’t be analyzed Protocol statistical section Analysis populations, primary and key secondary analyses, multiplicity strategy Vague protocol language the SAP can’t properly fix later Interim analysis plan Timing, stopping boundaries, firewall and DMC charter input Type I error inflation; questions about trial integrity eCRF review Confirming every variable needed for analysis is collected in analyzable form Missing dates, free text where codes were needed SAP draft and mock shells Early analysis plan and table layouts Programming starts late; database lock slips Data standards plan Mapping approach to CDISC SDTM and ADaM Costly retrospective conversion before submission
| Pre-study deliverable | What the partner owns | What goes wrong if it’s late |
| Endpoint and estimand definition | Precise definition of what is measured, when, and how intercurrent events are handled | Primary analysis can’t be computed as intended; mid-study amendments |
| Sample size and power | Assumptions, justification, sensitivity to those assumptions | Underpowered trial, or an oversized and overpriced one |
| Randomization and stratification | Scheme, block sizes, stratification factors, specifications for the IxRS | Imbalance, or strata that can’t be analyzed |
| Protocol statistical section | Analysis populations, primary and key secondary analyses, multiplicity strategy | Vague protocol language the SAP can’t properly fix later |
| Interim analysis plan | Timing, stopping boundaries, firewall and DMC charter input | Type I error inflation; questions about trial integrity |
| eCRF review | Confirming every variable needed for analysis is collected in analyzable form | Missing dates, free text where codes were needed |
| SAP draft and mock shells | Early analysis plan and table layouts | Programming starts late; database lock slips |
| Data standards plan | Mapping approach to CDISC SDTM and ADaM | Costly retrospective conversion before submission |
A few of these deserve a closer look.
Clinical Trial Design Statistics: Endpoints, Estimands and Sample Size
Endpoint selection and estimands come first
The ICH E9(R1) addendum on estimands, finalized by FDA in May 2021, asks sponsors to define precisely what treatment effect the trial is estimating. That includes how intercurrent events, such as treatment discontinuation or rescue medication, are handled.
This is a design decision. If the estimand says rescue medication use is handled by a specific strategy, the eCRF has to capture rescue medication with dates and times. If it doesn’t, the estimand is just wishful thinking.
Study design statistical considerations that usually need input before protocol sign-off include:
- Choice of primary endpoint and the timepoint it’s assessed at
- Handling of intercurrent events under the estimand framework
- Multiplicity across endpoints, doses or subgroups
- Whether an adaptive element is justified. FDA’s guidance on adaptive designs puts heavy weight on prespecification, which only works if the statistician is at the table during design.
Sample size and power calculation for clinical trials
A clinical trial sample size determination rests on a handful of assumptions: the size of the treatment effect you expect, how variable the outcome is, the significance level, the power you want, and how many participants you expect to drop out. Small changes in any of them can move the required sample size a lot.
That’s why the justification matters more than the number. A good statistical power calculation comes with the source for each assumption and a sensitivity analysis showing what happens if the effect is smaller or the variability larger. A single number with nothing behind it is a warning sign.
Statistical Analysis Plan Development: Start Early, Finish Before Unblinding
ICH E9 allows the statistical analysis plan to be written as a separate document, completed after the protocol is final. It also expects the SAP to be finalized before the blind is broken, with formal records of when each happened (Section 5.1).
That leaves room for a late SAP. Room isn’t the same as advice. An SAP drafted early, alongside the protocol or shortly after, does three useful things:
- It exposes ambiguities in the protocol while they can still be fixed cheaply.
- It tells data management exactly which variables and derivations the analysis depends on.
- It gives SAS programming for the clinical trial a specification to build against from the first data transfer.
The SAP will be revised. Blind data review usually prompts changes. What matters is that version one exists before the data do.
How to develop a statistical analysis plan
There’s no single template, but SAP development for clinical trials usually follows the same path:
- Start from the statistical section of the protocol and expand it, without contradicting it.
- Define the analysis populations and how each participant is assigned to them.
- Specify the primary and key secondary analyses, including the handling of missing data and intercurrent events.
- Set out the multiplicity approach and any interim analysis rules.
- Draft mock table shells so everyone can see what the outputs will look like.
- Have data management and programming review it, then version it and keep a record of every change.
A statistician reviewing the eCRF asks different questions from a data manager. The data manager asks whether a field can be completed and cleaned. The statistician asks whether the collected value can produce the planned analysis. You need both, and it’s why our clinical data management services and biostatistics work are set up to talk to each other early.
Things a statistical eCRF review commonly catches:
- Dates collected without times, when timing relative to dosing matters
- Free-text fields where coded responses were needed for analysis
- Missing reason-for-discontinuation categories that the estimand strategy depends on
If you want to see how the build side works, our guide to EDC data management services explains how the eCRF, validation and queries fit together.
Standards planning belongs here too. FDA requires standardized study data for NDAs, BLAs, ANDAs and commercial INDs, and lists CDISC SDTM for clinical data and ADaM for analysis data among its supported standards. Designing collection with SDTM mapping and ADaM derivations in mind avoids the expensive alternative: converting legacy-format data after the fact, under submission pressure. Your statistical programming team and your regulatory affairs lead should both have a view on this before the eCRF is final.
A Pre-Study Timeline for Statistical Deliverables
Biostatistical consulting is easiest to hold to account when each deliverable is tied to a study milestone. The exact sequence varies with phase and design, but a workable pattern looks like this:
| Study milestone | Statistical deliverable due | Who signs off |
| Protocol synopsis | Endpoint and estimand proposal; preliminary sample size | Sponsor clinical lead, statistician |
| Final protocol | Statistical section, sample size justification, multiplicity approach | Sponsor, statistician, medical lead |
| Randomization setup | Randomization specification and IxRS requirements | Statistician, clinical operations |
| eCRF build | Statistical eCRF review comments resolved | Data management, statistician |
| DMC charter (if applicable) | Interim analysis plan and firewall procedures | DMC, sponsor, independent statistician |
| First patient in | SAP draft version 1 and mock table shells | Sponsor, statistician |
Two rows get skipped most often: the eCRF review and the first SAP draft. They’re also the two that decide how smoothly database lock goes. If your partner can’t tell you which of these they own, and by when, ask before the contract is signed.
How Early Biostatistics Involvement Affects Database Lock Timelines
Most of the time lost at database lock was lost months earlier. Early clinical trial biostatistics support moves work out of the closeout window.
| Activity | Statistics involved from design | Statistics brought in after enrollment |
| SAP | Drafted early; refined during the study | Written under time pressure near lock |
| Table shells | Agreed before data arrive | Negotiated during closeout |
| ADaM specifications | Built iteratively against real transfers | Built in a rush on final data |
| Programming | Validated on interim transfers; dry runs before lock | Starts late; first runs on locked data |
| Data issues affecting analysis | Flagged as queries during the study | Discovered at blind review, when fixing is hardest |
The pattern is consistent. Analysis-driven data problems surface either during the study, as ordinary queries, or at lock, as delays.
Biostatistical Considerations by Trial Phase
The ownership map applies to every phase, but the weight shifts as a program matures.
| Phase | Where statistical effort concentrates before first patient in |
| Phase I | Dose-escalation rules, cohort decision criteria, pharmacokinetic analysis planning, safety review triggers |
| Phase II | Dose selection, go/no-go criteria, endpoint sensitivity, sample size under uncertain effect sizes |
| Phase III | Confirmatory hypotheses, estimands, multiplicity control, interim analyses, full SAP rigor |
| Post-approval | Real-world or registry comparisons, pre-specified subgroup work, safety signal methods |
A Phase I study with sparse data still needs someone to define the escalation rules precisely before dosing starts. A Phase III study needs everything on the list, in writing, before the protocol goes to regulators. One pattern holds across phases: the earlier a statistician tests the assumptions, the fewer surprises the data deliver.
Phase II is where that pays off most visibly. Its results set the effect size assumptions for Phase III, so a loosely planned Phase II analysis produces a loosely justified Phase III sample size. Asking your partner to show how the Phase II analysis will feed the next study’s design is a fair test of whether they’re thinking one step ahead. It also tells you whether the same team can carry the program forward. Continuity across phases means the person defending the Phase III sample size already knows why the Phase II endpoints were chosen, and which assumptions were weakest.
When Should You Involve a Biostatistician in a Clinical Trial?
Ideally at the protocol synopsis stage, and certainly before the statistical section is final. A late start doesn’t just slow the analysis. It narrows your options:
- Amendments. Fixing an endpoint definition or visit window mid-study means a protocol amendment, ethics and regulatory review, and site retraining.
- Unrecoverable data. A variable that was never collected can’t be cleaned into existence.
- Weaker analyses. When the data don’t support the planned approach, the fallback is often a less efficient method, or more sensitivity analyses to defend it.
- Regulatory questions. Analysis decisions made or changed close to unblinding draw scrutiny, because ICH E9 expects the SAP to be final before the blind is broken.
A few of these can be absorbed. All of them together can cost a trial its credibility.
Choosing a Biostatistics Consulting Partner: What to Ask Before You Sign
Outsourcing decisions are often made on price and headcount. The better signal is how a partner talks about pre-study work.
| Question to ask | A strong answer includes | A warning sign |
| When would your statistician first review our protocol? | Before final protocol sign-off | “Once the SAP is due” |
| Who writes the sample size justification, and what assumptions does it rest on? | Named statistician; assumptions sourced and stress-tested | A single number with no sensitivity analysis |
| Do you review the eCRF against the planned analysis? | Yes, with documented comments | Only data management reviews the eCRF |
| How do you handle estimands and intercurrent events? | Specific references to ICH E9(R1) strategies | General answers about “missing data” |
| When do you produce the SAP and table shells? | Draft early; shells agreed well before lock | “After last patient in” |
| How are SDTM and ADaM built and validated? | Independent double programming or equivalent QC; traceability from SDTM to ADaM to outputs | No clear validation approach |
| Who is the named statistician, and will they stay on the study? | A named lead with continuity plans | Rotating resources |
A statistical consulting partner that answers these well will usually push back on your protocol too. That pushback is part of the value.
At Weltrix, our biostatistics consulting services start at the design stage, with study design and protocol input, and carry through to statistical analysis and regulatory support. The aim is for your biostatistics team to work as an extension of yours, from the first protocol draft to the final tables.Frequently Asked Questions
Q. What are biostatistics consulting services?
Biostatistics consulting services cover the statistical work around a clinical trial: study design, endpoint and estimand definition, sample size and power calculation, randomization, the statistical analysis plan, interim analyses, data standards planning and the final analysis. The most valuable part happens before enrollment.
Q. What should a biostatistics partner be involved in before the trial begins?
A biostatistics partner should be involved in endpoint and estimand definition, sample size calculation, randomization design, the protocol’s statistical section, interim analysis planning, eCRF review, the first SAP draft and data standards planning for CDISC SDTM and ADaM.
Q. What does a biostatistician do before a clinical trial starts?
They help define the question the trial answers. That means choosing the endpoints and estimands, calculating sample size and power, designing randomization, writing the statistical section of the protocol, reviewing the eCRF against the planned analysis, and drafting the SAP.
Q. When should you involve a biostatistician in a clinical trial?
As early as the protocol synopsis, and definitely before the protocol’s statistical section is final. Bringing one in after enrollment starts limits what can be fixed without an amendment.
Q. What is a statistical analysis plan (SAP), and why does it matter this early?
A statistical analysis plan is the detailed document describing how trial data will be analyzed, expanding on the statistical section of the protocol. ICH E9 expects it to be finalized before unblinding. Drafting it early matters because it exposes protocol ambiguities while they’re cheap to fix and tells data management and programmers what the analysis needs.
Q. How does early biostatistics involvement affect database lock timelines?
Early involvement shortens the path to lock. Table shells, ADaM specifications and programs are built and tested during the study, and analysis-critical data issues are raised as ordinary queries instead of surfacing at blind review.
Q. What’s the risk of bringing in biostatistics support too late?
The main risks are protocol amendments to fix endpoints or visit windows, data that was never collected and can’t be recovered, weaker fallback analyses, and regulatory questions about analysis decisions made close to unblinding. Some of these can be absorbed. Unrecoverable data can’t.
Q. What should sponsors ask when choosing a biostatistics CRO partner?
Ask when the statistician first reviews the protocol, who owns the sample size justification, whether the eCRF is reviewed against the planned analysis, how estimands are handled under ICH E9(R1), when the SAP and shells are produced, how SDTM and ADaM outputs are validated, and whether the named statistician stays on the study.
Conclusion
The value of biostatistics consulting services is decided before the first patient signs consent. By the time data arrive, the endpoint, the estimand, the sample size and the fields on the eCRF are fixed, and the analysis can only be as good as those choices. Weltrix brings biostatistics into protocol development for exactly that reason. Whichever partner you choose, ask one thing first: which page of your protocol will their statistician read before you sign it?
Key Takeaways
- ICH E9 expects the protocol’s statistical section to describe the principal features of the analysis, so statistical input has to come before protocol sign-off.
- Estimands under ICH E9(R1) are design decisions. The eCRF must capture the data each estimand strategy needs.
- A sample size is only as good as its assumptions. Ask for the sources and a sensitivity analysis.
- The SAP can be finalized later, but it must be final before unblinding, and an early draft exposes protocol gaps cheaply.
- A statistical eCRF review catches missing dates, times and codes that data cleaning can never recover.
- FDA requires standardized study data for NDAs, BLAs, ANDAs and commercial INDs. Planning SDTM and ADaM at setup avoids retrospective conversion.
- Early statistics involvement moves programming and data issues out of the database lock window.


Leave A Comment