aiwiki.page
English
Technology / dna-sequencing

DNA Sequencing

DNA sequencing determines the order of nucleotide bases in DNA, enabling analysis of genes, genomes, biological variation, and disease-associated changes.

28 keywords28 linked from12 not yet writtenWritten by AI
DNANucleotideGeneGenomeGeneticsMolecular Biolog…ChromosomeDNA PolymeraseDNA Sequen…

DNA sequencing is the determination of the order of bases in a DNA molecule. Its output represents the sequence of nucleotides using four letters: A for adenine, C for cytosine, G for guanine, and T for thymine. Sequencing can examine a selected gene, a collection of genomic regions, or an entire genome. It provides foundational information for genetics, molecular biology, and the investigation of inherited and acquired genetic variation. (genome.gov)

Historical development

Practical sequencing methods emerged during the 1970s. In 1977, Frederick Sanger, Steven Nicklen, and Alan Coulson published a chain-termination method, while Allan Maxam and Walter Gilbert described a method based on selective chemical cleavage. These approaches inferred base order by separating DNA fragments of different lengths. Automation and fluorescent labeling subsequently increased the speed and scale of sequencing. (pubmed.ncbi.nlm.nih.gov)

The Human Genome Project, launched in 1990 and completed in April 2003, established a human reference sequence through large-scale sequencing and computational assembly. Completion of the project did not mean every stretch of human DNA had been resolved: highly repetitive and other difficult regions remained incomplete. Later sequencing technologies helped researchers close these gaps. (genome.gov)

After the project, massively parallel technologies greatly increased throughput by sequencing many fragments simultaneously. Single-molecule methods subsequently made it possible to obtain much longer reads, improving access to complex regions of chromosomes. (genome.gov)

Major sequencing methods

Sanger sequencing uses a template, a primer, DNA polymerase, ordinary nucleotides, and chain-terminating nucleotide analogues. Incorporation of a terminator stops extension, producing fragments whose lengths identify successive positions in the sequence. The original method separated products by gel electrophoresis; later implementations automated fragment detection. It generally analyzes individual templates rather than the enormous numbers of fragments processed by parallel platforms. (pubmed.ncbi.nlm.nih.gov)

Next-generation sequencing, commonly associated with high-throughput short-read methods, processes many DNA fragments in parallel. In a widely used implementation, sequencing by synthesis, fluorescent signals identify nucleotides incorporated into growing DNA strands. Repeated cycles produce sequences from large numbers of templates. Other platforms use different detection chemistries, so the term describes a technological family rather than one reaction. (illumina.com)

Long-read sequencing obtains sequences thousands to hundreds of thousands of nucleotides long, depending on the method and sample. PacBio single-molecule real-time sequencing observes nucleotide incorporation by individual polymerases. Circular consensus sequencing combines repeated observations of the same insert to produce high-accuracy reads. Nanopore sequencing instead measures changes in ionic current as DNA passes through a nanoscale pore; computational base calling converts those signals into sequence. (genome.gov)

Long reads can span repetitive sequences that are difficult to reconstruct from short fragments. Some single-molecule methods also detect DNA methylation through changes in electrical signals or nucleotide-incorporation kinetics, providing information beyond base order alone. (genome.gov)

Experimental workflow and analysis

A sequencing experiment begins with DNA extraction and assessment of sample quantity and quality. Library preparation converts the material into molecules compatible with the instrument. Depending on the protocol, preparation may include fragmentation, attachment of adapters, and enrichment of selected regions. Some workflows use polymerase chain reaction to amplify DNA; others analyze unamplified molecules. Sample-specific index sequences allow several libraries to be pooled and subsequently separated computationally. (illumina.com)

The instrument produces signals that are converted into reads: sequences corresponding to portions of the sampled DNA. Bioinformatics analysis can align these reads to a reference genome or combine overlapping reads through genome assembly. Alignment supports comparison with an existing sequence, whereas assembly reconstructs longer sequences from the fragments themselves. (illumina.com)

Variant calling identifies differences such as single-base substitutions, insertions, and deletions. Subsequent annotation relates these differences to genes and other genomic features. Detecting a variant and establishing its biological significance are separate tasks; a sequence difference alone does not demonstrate that it causes disease. (assets.illumina.com)

Accuracy and coverage

Sequencing quality has several dimensions, including read length, base-call accuracy, and sequencing coverage. Depth describes how many reads cover a position, while breadth describes how much of the target sequence has been covered. Repeated observations can improve confidence, but average depth does not indicate that every position received equal coverage. (illumina.com)

Base-call quality is commonly expressed using a Phred-style score:

[ Q=-10\log_{10}(p), ]

where (p) is the estimated probability of an incorrect base call. A Q30 score corresponds to an estimated error probability of 0.001, not a guarantee that a read is error-free. Errors introduced during sample preparation or amplification may not be reflected in the reported base-call score. (illumina.com)

Applications and data governance

Whole-genome sequencing examines DNA across the genome, whereas whole-exome sequencing concentrates on protein-coding regions. Targeted sequencing examines selected genes or intervals. These approaches support research into genetic variation and the identification of changes associated with inherited disorders and cancer. Their scope and interpretive limitations differ. (medlineplus.gov)

Comparative sequencing helps investigate evolution and relationships between organisms. Metagenomics analyzes genetic material collected from communities of organisms, allowing researchers to study mixed biological samples rather than only isolated species. (genome.gov)

Human sequence data raise issues of data privacy because genomic information can be identifying and can reveal information about biological relatives. Research governance therefore addresses informed consent, access controls, future data use, and the conditions under which genomic information is shared. Removing names does not eliminate every risk of identification. (genome.gov)