Clinical trial data management from first patient to database lock: what the data management plan has to settle before enrolment, how query volume is designed down rather than worked through, and the standards a submission dataset has to arrive in
Data management is judged on one day, the day of database lock, and everything that makes that day easy was decided before the first patient was enrolled. Most of the pain in a trial's data comes from a case report form that asked an ambiguous question and an edit check nobody wrote. This page covers the decisions in the order they have to be made.
- the electronic records and signatures rule the data system meets
- Part 11
- the investigational new drug regulation the trial runs under
- Part 312
- the standard family submission datasets are prepared in
- CDISC
Figures in this panel are the regulations and the data standard a clinical trial's data is collected and submitted under, named from the regulations themselves and linked in the sources below. They are identifiers, not prices: BioBricks publishes verified prices for synthesis services only, and does not imply a software price index it has not measured.
- 4 vendor service pages verifiedevery figure matched verbatim to the vendor's page
- Quoted and dated, never estimatedlast verification pass 2026-08-24
- 1 service classes coveredeach with measured search demand behind it
From plan to lock
- Write the data management plan before enrolment. It names the systems, the roles, the coding dictionaries and their versions, the edit check specification, the query workflow and the lock criteria. Written after enrolment starts, it documents what happened rather than governing it.
- Design the case report form to prevent queries, not to collect them. Every ambiguous field becomes a query, and every query costs a site coordinator and a monitor time. Constrained lists instead of free text, one question per field, and a pilot with real site staff before go-live remove most of the query volume before it exists.
- Build edit checks that fire at entry. A range check, a consistency check between related fields and a required field rule catch the error in front of the person who knows the answer. A check that runs in a nightly batch produces a query a week later, to somebody who has to find the source document again.
- Code adverse events and medications with a stated dictionary version. Coding to a named version, with a documented process for terms the dictionary does not carry, is what makes the safety analysis defensible. Upgrading a dictionary mid-study is a decision with consequences and needs a documented plan.
- Produce submission-ready datasets in the standard from the start. Mapping to the standard tabulation model and the analysis datasets built on it is far cheaper designed in than retrofitted. A study that collects data in its own shape and maps at the end pays for that twice, in effort and in review questions.
External data is where reconciliation fails
Central laboratory results, imaging reads, devices and drug accountability arrive from other systems on other schedules with their own identifiers. Agree the transfer specification, the frequency and the reconciliation rules before the first transfer rather than after the first mismatch.
Reconcile continuously rather than at lock. A subject identifier mismatch found in month two is an email; the same mismatch found at lock is a delay everybody sees.
Validation and the audit trail
The data system has to be validated for its intended use, with the audit trail on and reviewed rather than merely enabled. Audit trail review is an expectation, and a study that never looked at its own audit trail cannot say what it shows.
Keep the validation documentation current through upgrades. A system validated at go-live and upgraded twice since is, on paper, unvalidated.
Working with a provider
Where data management is outsourced, the sponsor still owns the obligation. Agree who writes the plan, who builds the checks, who closes queries, what the escalation path is and what the data looks like when it is handed back.
Ask for the data in an analysable, standard form at agreed intervals rather than only at the end. A sponsor seeing its own data monthly finds the design problems while they can still be fixed.
Common questions
- What does a data management plan have to cover?
- Systems and their validation, roles and access, the case report form and edit check specification, coding dictionaries and versions, the query workflow and timelines, external data handling, and the criteria for database lock.
- How do I reduce query volume?
- At form design. Constrained fields rather than free text, one question per field, edit checks that fire at entry, and a pilot with the people who will actually fill it in. Queries are a symptom of a form that asked badly.
- Why do submission standards matter before analysis?
- Because retrofitting them is expensive and error-prone. Collecting and mapping into the standard tabulation model as the study runs means the analysis datasets and the submission package are a continuation rather than a project.
- What is required to lock a database?
- All data entered, all queries closed, all external data reconciled, coding complete and quality-reviewed, and the audit trail intact. The criteria belong in the plan so lock is a checklist rather than a negotiation.
Get a shortlist for your project
Browse by service class
Sources
Cite or embed this figure
The median advertised gene synthesis price per base pair in the US research synthesis services market was $0.11 in August 2026, across 4 verified vendor service pages recorded in BioBricks Synthesis Price Index.
Cite as: "BioBricks Synthesis Price Index", updated 2026-08-24, https://biobricks.org/clinical-trial-data-management/.