Clinical trial data management from first patient to database lock: what the data management plan has to settle before enrolment, how query volume is designed down rather than worked through, and the standards a submission dataset has to arrive in
Data management is judged on one day, the day of database lock, and everything that makes that day easy was decided before the first patient was enrolled. Most of the pain in a trial's data comes from a case report form that asked an ambiguous question and an edit check nobody wrote. This page covers the decisions in the order they have to be made.
- the electronic records and signatures rule the data system meets
- Part 11
- the investigational new drug regulation the trial runs under
- Part 312
- the standard family submission datasets are prepared in
- CDISC
Figures in this panel are the regulations and the data standard a clinical trial's data is collected and submitted under, named from the regulations themselves and linked in the sources below. They are identifiers, not prices: BioBricks publishes verified prices for synthesis services only, and does not imply a software price index it has not measured.
- 4 vendor service pages verifiedevery figure matched verbatim to the vendor's page
- Quoted and dated, never estimatedlast verification pass 2026-08-24
- 1 service classes coveredeach with measured search demand behind it
From plan to lock
- Write the data management plan before enrolment. It names the systems, the roles, the coding dictionaries and their versions, the edit check specification, the query workflow and the lock criteria. Written after enrolment starts, it documents what happened rather than governing it.
- Design the case report form to prevent queries, not to collect them. Every ambiguous field becomes a query, and every query costs a site coordinator and a monitor time. Constrained lists instead of free text, one question per field, and a pilot with real site staff before go-live remove most of the query volume before it exists.
- Build edit checks that fire at entry. A range check, a consistency check between related fields and a required field rule catch the error in front of the person who knows the answer. A check that runs in a nightly batch produces a query a week later, to somebody who has to find the source document again.
- Code adverse events and medications with a stated dictionary version. Coding to a named version, with a documented process for terms the dictionary does not carry, is what makes the safety analysis defensible. Upgrading a dictionary mid-study is a decision with consequences and needs a documented plan.
- Produce submission-ready datasets in the standard from the start. Mapping to the standard tabulation model and the analysis datasets built on it is far cheaper designed in than retrofitted. A study that collects data in its own shape and maps at the end pays for that twice, in effort and in review questions.
External data is where reconciliation fails
Central laboratory results, imaging reads, devices and drug accountability arrive from other systems on other schedules with their own identifiers. Agree the transfer specification, the frequency and the reconciliation rules before the first transfer rather than after the first mismatch.
Reconcile continuously rather than at lock. A subject identifier mismatch found in month two is an email; the same mismatch found at lock is a delay everybody sees.
Validation and the audit trail
The data system has to be validated for its intended use, with the audit trail on and reviewed rather than merely enabled. Audit trail review is an expectation, and a study that never looked at its own audit trail cannot say what it shows.
Keep the validation documentation current through upgrades. A system validated at go-live and upgraded twice since is, on paper, unvalidated.
Working with a provider
Where data management is outsourced, the sponsor still owns the obligation. Agree who writes the plan, who builds the checks, who closes queries, what the escalation path is and what the data looks like when it is handed back.
Ask for the data in an analysable, standard form at agreed intervals rather than only at the end. A sponsor seeing its own data monthly finds the design problems while they can still be fixed.
A clinical trial edc, and where the laboratory data meets it
Electronic data capture holds what the site records about a subject, and laboratory results usually arrive from somewhere else entirely: an instrument, a laboratory information system, or a central laboratory's transfer file. The interesting part is the boundary rather than either system, because that is where units, reference ranges, subject identifiers and repeat visits get reconciled.
Agree the transfer specification before the first sample: the identifier that joins the two, the units and their precision, how a repeat or a corrected result is represented, and who reconciles a mismatch. A trial that discovers at analysis that two systems disagree about a subject identifier spends months on it.
A data integrity risk assessment, and what it examines
A data integrity assessment walks the life of a record rather than auditing a system: where data is created, whether it is attributable and timestamped, whether the original is preserved when it is transcribed, who can change it and whether the change is visible, how it is backed up and how it is retrieved years later. The findings that matter are usually mundane, a shared login, a spreadsheet between an instrument and a database, an audit trail nobody reviews, a clock nobody synchronises. Score by patient and product impact, and fix the paper-to-system boundaries first.
gxp data integrity and what a system has to show
gxp data integrity means the record is attributable, legible, contemporaneous, original and accurate, and a system demonstrates it with unshared accounts, an audit trail that cannot be disabled, time synchronisation and a backup that has been restored at least once. Most findings in this area are about the process around the software rather than the software itself.
crispr data analysis and the pipeline that has to be recorded
crispr data analysis means calling indels from amplicon reads, quantifying the spectrum, and assessing off target sites, and the pipeline's version and parameters are part of the result because different callers disagree on the same reads. A frame preserving indel can leave a functional protein, which is why the spectrum rather than a single efficiency number is reported.
crispr controls and the set a claim needs
crispr controls are a non targeting guide at the same delivery and selection, at least two independent guides against the target, an unedited population carried in parallel, and where the phenotype matters a rescue. Without the second guide an off target effect and the intended one cannot be separated, which is the commonest gap in a published figure.
Common questions
- What does a data management plan have to cover?
- Systems and their validation, roles and access, the case report form and edit check specification, coding dictionaries and versions, the query workflow and timelines, external data handling, and the criteria for database lock.
- How do I reduce query volume?
- At form design. Constrained fields rather than free text, one question per field, edit checks that fire at entry, and a pilot with the people who will actually fill it in. Queries are a symptom of a form that asked badly.
- Why do submission standards matter before analysis?
- Because retrofitting them is expensive and error-prone. Collecting and mapping into the standard tabulation model as the study runs means the analysis datasets and the submission package are a continuation rather than a project.
- What is required to lock a database?
- All data entered, all queries closed, all external data reconciled, coding complete and quality-reviewed, and the audit trail intact. The criteria belong in the plan so lock is a checklist rather than a negotiation.
Get a shortlist for your project
Browse by service class
Sources
Cite or embed this figure
The median advertised gene synthesis price per base pair in the US research synthesis services market was $0.11 in August 2026, across 4 verified vendor service pages recorded in BioBricks Synthesis Price Index.
Cite as: "BioBricks Synthesis Price Index", updated 2026-08-24, https://biobricks.org/clinical-trial-data-management/.