rna seq analysis: the design decisions that decide whether the data can answer the question, and which parts are worth outsourcing
Most disappointing transcriptome experiments were decided before any sequencing happened, in the choice of replicates, library type and how samples were distributed across batches. The analysis cannot rescue a design that confounded the comparison, and no depth substitutes for replication. This page covers the decisions that matter, in the order they are made, and which parts of the work are worth handing to somebody else.
- the NIH sequence archive the reference annotations descend from
- GenBank
- the NCBI similarity service pipelines call for identity checks
- BLAST
- the European service mirroring the public archives and their APIs
- EMBL-EBI
Names in this panel are the public archives and services an analysis actually depends on, linked in the sources below. They are identifiers, not prices: BioBricks publishes verified prices for synthesis services only, and does not imply a bioinformatics price index it has not measured.
- 4 vendor service pages verifiedevery figure matched verbatim to the vendor's page
- Quoted and dated, never estimatedlast verification pass 2026-08-24
- 1 service classes coveredeach with measured search demand behind it
Decisions in the order they are made
- Replicates beat depth. Statistical power for differential expression comes from biological replication far more than from reads per sample, and technical replicates measure the machine rather than the biology. When budget is fixed, more biological replicates at moderate depth almost always beats fewer at high depth.
- Library type follows the question. Poly-A selection suits intact eukaryotic messenger RNA; ribosomal depletion suits degraded, bacterial or non-polyadenylated transcripts. Stranded libraries resolve overlapping features and antisense transcription and are now the sensible default. These choices are made before sequencing and cannot be changed after.
- Randomise across batches. If every treated sample is prepared on one day and every control on another, treatment and batch are confounded and no analysis can separate them. Distribute groups across preparation batches and sequencing runs, and record which sample went where.
- Reference and annotation are a choice. The genome build and annotation release determine what can be quantified, and mixing releases across a project produces results that cannot be compared. Fix both at the start, record them, and use them throughout.
- Normalisation and the statistical method. Counts must be normalised for depth and composition before comparison, and the established differential expression packages handle that along with dispersion estimation. Spreadsheet fold-changes on raw counts are not an analysis, whatever the numbers look like.
What to outsource and what to keep
Library preparation and sequencing are commodity services and are usually best bought. Primary processing to a count matrix is routine and can be bought or run with a workflow manager. The biological interpretation is where the value is and it should stay with the people who know the system.
If you buy the analysis, require the count matrix, the software versions and parameters, and the quality control report, not only a list of differentially expressed genes. Without those, the result cannot be reproduced or re-examined.
Quality control that is worth reading
Read duplication, ribosomal content, mapping rate, gene body coverage and the clustering of samples before any statistical test. A sample that clusters with the wrong group in an unsupervised view is telling you something the differential test will obscure.
Keep the raw data and deposit it in a public archive when the work is published. Journals and funders increasingly require it, and preparing the submission at the end is far harder than keeping the metadata as you go.
Common questions
- How many replicates do I need for RNA seq?
- As many biological replicates as the budget allows, because power for differential expression comes from replication far more than from sequencing depth. Technical replicates measure the instrument rather than the biology.
- Poly-A selection or ribosomal depletion?
- Poly-A selection for intact eukaryotic messenger RNA; ribosomal depletion for degraded material, bacterial RNA or non-polyadenylated transcripts. The choice is made before sequencing and cannot be revisited.
- What should an outsourced analysis deliver?
- The count matrix, the software versions and parameters, the quality control report and the raw data, alongside the differential expression results. Without those the analysis cannot be reproduced.
- How do I avoid batch effects?
- Distribute treatment groups across preparation batches and sequencing runs rather than processing each group together, and record which sample went where. A confounded design cannot be fixed analytically.
Get a shortlist for your project
Browse by service class
Sources
Cite or embed this figure
The median advertised gene synthesis price per base pair in the US research synthesis services market was $0.11 in August 2026, across 4 verified vendor service pages recorded in BioBricks Synthesis Price Index.
Cite as: "BioBricks Synthesis Price Index", updated 2026-08-24, https://biobricks.org/rna-seq-analysis/.