Next-generation sequencing (NGS) takes a DNA or RNA sample through three broad stages — wet-lab library preparation, sequencing on the instrument, and computational analysis — to go from a tube of nucleic acid to an annotated list of genetic variants. Each stage has its own quality checkpoints, and a failure at any one of them quietly degrades everything downstream.

Stage 1: Wet-Lab Preparation

Extraction and Input QC

Genomic DNA is isolated by column or magnetic bead chemistry, with typical inputs of 100–500 ng for standard workflows and as little as 1–10 ng for PCR-based panels. RNA workflows add an extra step — rRNA depletion or polyA selection, followed by reverse transcription to cDNA — before library construction can begin. Every sample is quantified (Qubit dsDNA HS), checked for purity (A260/280 of 1.8–2.0), and checked for integrity (DIN ≥ 7 for DNA, RIN ≥ 8 for RNA) before it proceeds.

Library Preparation

The genome is sheared into fragments of roughly 350–550 bp — by sonication, enzymatic fragmentation, or tagmentation — then made sequenceable through end repair, dA-tailing, and adapter ligation. Unique dual indexes (8–10 bp barcodes on each end) are attached during PCR so that 96 or more samples can be pooled and sequenced together on a single flow cell without their reads getting mixed up.

Enrichment and Pooling

For targeted applications, hybrid capture with biotinylated probes pulls down only the regions of interest, trading breadth of coverage for depth. Typical depth targets vary widely by application: whole genome sequencing runs at 30–40×, exome sequencing at roughly 100×, and sensitive somatic panels at 500–1000× to catch low-frequency variants.

Stage 2: Sequencing on the Instrument

Sequencing by synthesis happens in three steps. Denatured single strands hybridize to a lawn of complementary oligos on the flow cell; bridge amplification (or an equivalent chemistry) copies each molecule roughly a thousand times in place, producing a monoclonal cluster of identical strands. The instrument then cycles through incorporation of one fluorescently labeled, chain-terminating nucleotide at a time, imaging the flow cell after each cycle before cleaving the dye and blocker to allow the next base — typically 150 cycles per read. In a paired-end run, the process repeats from the opposite end of each fragment after an index read, so both ends of the same molecule are captured.

MetricTypical Target
Read configuration2 × 150 bp paired-end
Clusters per flow cell~400M–20B
Bases ≥ Q30> 90%
Mapping rate> 95%
Duplicate rate< 15%

Stage 3: The Bioinformatics Pipeline

Raw signal (BCL) is demultiplexed into per-sample FASTQ files by index, with adapters and low-quality read tails trimmed. Reads are then aligned to a reference genome, sorted, and duplicate reads are marked (or collapsed using unique molecular identifiers) before base quality recalibration. Variant callers identify SNVs and indels per sample, and structural variant and copy-number callers run in parallel. The resulting calls are filtered, annotated against population and clinical databases, and classified for review — moving from a raw gVCF to an annotated report a scientist or clinician can act on.

Where NGS Runs Commonly Go Wrong

  • PCR duplicates: low input or over-cycled library prep inflates apparent coverage without adding new information
  • Index hopping: free adapters on patterned flow cells cross-contaminate samples — unique dual indexing is the fix
  • GC bias: extreme GC-content regions lose coverage during amplification and cycling
  • Mismapping: short reads struggle to resolve repeats, HLA regions, and segmental duplications
  • Cross-sample contamination: caught by fingerprint concordance checks between samples

Frequently Asked Questions

How long does a full NGS workflow take?

Wet-lab preparation typically takes 1–2 days, sequencing runs 12–48 hours depending on the platform and read length, and bioinformatics analysis can take from a few hours to several days depending on depth and the number of samples.

What sequencing depth do I need?

It depends on the application — roughly 30–40× for whole genome sequencing, 100× for exome sequencing, and 500–1000× for sensitive somatic variant panels.

Why does input DNA quality matter so much?

Degraded or impure DNA produces shorter, biased fragments during library prep, which lowers mapping rate and coverage uniformity — problems that are difficult to correct for computationally after the fact.

Conclusion

A reliable NGS result depends on tight control at every stage — input QC, library preparation, sequencing quality metrics, and a validated bioinformatics pipeline. Our sequencing services cover the full workflow, from sample extraction through variant annotation.