Designing a bulk rna seq experiment: choosing between an mrna isolation kit, an mrna kit route and a ribosomal rna depletion kit, setting a depth that answers the question, and why replicates buy more than reads do
Bulk RNA sequencing is cheap enough that the design decisions now cost more than the sequencing. The three that matter are how the library is made, how deep it is sequenced and how many biological replicates there are, and only the last of those is usually short. This page covers the choices in the order they have to be made.
- reads per sample commonly used for gene level differential expression
- 20-30M reads
- the authentication guidance a funded study is expected to follow
- NIH rigor
- the competence standard an accredited sequencing provider holds
- ISO 17025
Figures in this panel are the working depth convention for this experiment type and the guidance and accreditation a study and its provider sit under, linked in the sources below. They are identifiers, not prices: BioBricks publishes verified prices for synthesis services only, and does not imply a sequencing price index it has not measured.
- 4 vendor service pages verifiedevery figure matched verbatim to the vendor's page
- Quoted and dated, never estimatedlast verification pass 2026-08-24
- 1 service classes coveredeach with measured search demand behind it
The design decisions, in order
- Choose the library chemistry from the RNA you have. Poly-A selection with an mrna isolation kit captures intact polyadenylated transcripts and is the efficient choice for good quality eukaryotic RNA. A ribosomal rna depletion kit keeps non-polyadenylated and degraded material, which is what degraded or archival samples, bacterial RNA and non-coding work require.
- Measure RNA integrity before you commit. An integrity number from a capillary system, not a gel photograph. Degraded RNA is where poly-A selection fails quietly, giving a strong three prime bias that looks like biology in the coverage plot and is not.
- Set depth from the question. Differential expression of moderately expressed genes is generally served by something in the region of twenty to thirty million reads per sample. Isoform level work, low abundance transcripts and allele specific analysis need considerably more, and going deeper never substitutes for having too few samples.
- Spend the marginal dollar on replicates. Below about four or five biological replicates per group, power is dominated by biological variance rather than by counting noise, so an extra sample buys far more than an extra ten million reads. This is the most common and most expensive design mistake in the field.
- Decide stranded and read length deliberately. A stranded protocol should be the default: it resolves overlapping genes and antisense transcription for no extra cost. Paired end and longer reads help isoform assignment and are largely wasted on straightforward gene level counting.
What to agree with the provider before samples ship
The library chemistry by name, the target depth per sample, the read configuration, the demultiplexing and the file formats you will receive, and who keeps the raw data and for how long. Half of these are assumed rather than agreed and all of them matter later.
Agree the failure policy too: what happens if a library fails quality control, whether it is remade, and who pays. That conversation is much easier before the samples are in a courier.
Analysis decisions that belong in the design
The reference genome and annotation version, the quantification tool and the statistical model should be written down before the data arrives. Choosing them afterwards, with the results visible, is how an analysis stops being a test of anything.
Decide the multiple testing approach and the effect size you care about in advance. A gene list thresholded on significance alone will be dominated by highly expressed genes with small changes.
When bulk is the wrong tool
Bulk averages across whatever cells are in the sample, so a change confined to a minority population disappears into the mean. If the question is about composition or about a rare subset, single cell or sorted bulk is the honest answer.
Sorted bulk is often the sensible compromise: it keeps the depth and the statistics of bulk while removing the composition confound, at the cost of a sorting step and its own biases.
Common questions
- Poly-A selection or rRNA depletion for bulk rna seq?
- Poly-A with an mRNA isolation kit for intact eukaryotic RNA and gene level expression. rRNA depletion for degraded or archival samples, for bacterial RNA, and whenever non-polyadenylated transcripts are part of the question.
- How many reads per sample do I need?
- For differential expression of moderately expressed genes, roughly twenty to thirty million is a common working figure. Isoform level and low abundance work need more. Decide from the question and then check what the analysis actually requires.
- How many replicates?
- More than three if the effect is modest. Biological variance, not sequencing depth, sets the power in almost every bulk experiment, so the marginal sample is worth far more than the marginal read.
- Does batch matter if everything is sequenced together?
- Library preparation batch matters as much as sequencing batch, and so does the day of extraction. Randomise groups across every batch boundary and record them, so batch can be modelled rather than confounded.
Get a shortlist for your project
Browse by service class
Sources
Cite or embed this figure
The median advertised gene synthesis price per base pair in the US research synthesis services market was $0.11 in August 2026, across 4 verified vendor service pages recorded in BioBricks Synthesis Price Index.
Cite as: "BioBricks Synthesis Price Index", updated 2026-08-24, https://biobricks.org/bulk-rna-seq/.