Low pass whole genome sequencing: choosing depth for the question
Depth is the main cost lever in genome sequencing and the main determinant of what can be concluded. Shallow coverage across the whole genome answers questions about large scale structure cheaply; calling individual variants, and especially rare ones in a mixture, needs depth and molecular tagging.
- electronic records and signatures, the clause behind an analysis record
- Part 11
- good laboratory practice for nonclinical studies, 21 CFR
- Part 58
- the competence standard a testing laboratory is assessed against
- 17025
The figures in this panel are regulation and standard identifiers, named from the documents themselves and linked below. They are not prices: BioBricks publishes verified prices for synthesis services only, and does not imply a price index it has not measured.
- 4 vendor service pages verifiedevery figure matched verbatim to the vendor's page
- Quoted and dated, never estimatedlast verification pass 2026-08-24
- 1 service classes coveredeach with measured search demand behind it
Scoping a genome project
- Match depth to the smallest thing you must detect. Copy number and ancestry survive shallow coverage. Germline variant calling needs moderate depth. Rare somatic variants in a mixture need deep coverage plus molecular identifiers to separate signal from error.
- Judge library chemistry on evenness. Coverage uniformity and duplicate rate matter more than nominal yield, because uneven coverage leaves regions uncalled at any average depth.
- Use molecular identifiers where the fraction is small. Tagging original molecules distinguishes a true low frequency variant from an amplification or sequencing error. Without them, low frequency calling is guesswork.
- Keep a targeted method for targeted questions. When the variant is known and the question is present or absent, amplification is faster and cheaper than sequencing. Not everything needs a genome.
- Plan storage and analysis before generating data. Genome scale data is large and long lived. Where it lives, who pays after the project and how it is analysed reproducibly are decided at the start or not at all.
- Agree deliverables including the pipeline. Raw reads, alignments, variant calls with their filters, the reference and annotation versions, and runnable code. A variant table alone cannot be re-analysed.
Average depth hides the problem
A stated average coverage says nothing about the regions that received almost none, and those regions are systematically the same ones between samples. Uniformity, not average, is what determines callability.
Ask for per region coverage from validation samples, and check the regions your question depends on.
Data outlives the project
Genome scale data is generated in weeks and stored for years, and the storage cost accrues to whoever inherits it. Deciding where it lives and who pays is part of the project plan.
Keep the raw reads and the pipeline; intermediate files can usually be regenerated and are the bulk of the volume.
targeted methylation sequencing, and the chemistries behind it
Reading methylation at chosen loci combines an enrichment with a conversion chemistry. Bisulfite converts unmethylated cytosine to uracil and is the long-standing reference, at the cost of severe DNA damage and a reduced-complexity library; enzymatic conversion achieves the same read with far less fragmentation, which matters for low-input and cell free DNA. Enrichment is by capture probes, by amplicon panels designed for converted sequence, or by methyl-binding proteins. Specify the target list, the input amount and the conversion, since a methylation percentage is only comparable within one chemistry.
dna genotyping, and the methods by scale
Genotyping asks which alleles are present at known positions, and the method follows the count. A handful of variants is a targeted assay, allele-specific PCR or a probe-based call. Hundreds to a million is an array, which is what population and agricultural work runs on. A whole genome or exome answers the question and finds new variants at a higher price, and low pass sequencing with imputation against a reference panel has taken much of the array's territory where a suitable panel exists. Call rate and reproducibility are the metrics to compare.
Common questions
- What can shallow coverage answer?
- Copy number, large structural features and ancestry, at a small fraction of the cost. It cannot reliably call individual variants, and describing it as whole genome sequencing without the depth caveat overstates it.
- Why do rare variants need molecular identifiers?
- Because at low allele fraction the true signal is comparable to the error rate. Tagging original molecules lets errors introduced after tagging be identified and removed.
- Is sequencing always better than amplification?
- No. For a known variant with a yes or no answer, a targeted amplification assay is faster, cheaper and easier to validate. Sequencing earns its place when the question is open.
Get a shortlist for your project
Browse by service class
Sources
Cite or embed this figure
The median advertised gene synthesis price per base pair in the US research synthesis services market was $0.11 in August 2026, across 4 verified vendor service pages recorded in BioBricks Synthesis Price Index.
Cite as: "BioBricks Synthesis Price Index", updated 2026-08-24, https://biobricks.org/low-pass-whole-genome-sequencing/.