Phylogenetics is the branch of biology concerned with reconstructing and studying evolutionary relationships. It uses evidence from inherited characteristics, including anatomical features and molecular sequences, to infer histories of common ancestry. These histories, called phylogenies, are usually represented as trees, although networks may be needed when lineages exchange genetic material. Phylogenetics provides a historical framework for understanding evolution and contributes to biological classification, comparative research, and the interpretation of genetic data. An inferred phylogeny is a hypothesis supported by evidence and analytical assumptions, rather than a direct observation of the past. (pubmed.ncbi.nlm.nih.gov)
Scope and basic principles
Phylogenetic relationships concern shared ancestry, not simply overall resemblance. Two organisms can resemble one another because they inherited features from a common ancestor, but similarities can also arise independently. Conversely, closely related lineages may become markedly different. The central task is therefore to distinguish patterns of inheritance from other sources of similarity and to evaluate which evolutionary history best explains the evidence. (pubmed.ncbi.nlm.nih.gov)
A homology is a correspondence attributable to common ancestry. A shared derived character, or synapomorphy, is a homologous character state that originated in an ancestral lineage and helps identify its descendants. Shared ancestral states are less informative about relationships within the group being examined. Homoplasy—similarity arising through independent changes or evolutionary reversals—can obscure the historical signal. Cladistics formalized the use of shared derived characters to identify groups united by common descent. (pubmed.ncbi.nlm.nih.gov)
Phylogenetics and taxonomy overlap but are not identical. Taxonomy concerns the identification, naming, and classification of organisms; phylogenetics investigates their evolutionary relationships. Classifications may be revised to reflect phylogenetic evidence, but the assignment of names and taxonomic ranks involves conventions beyond tree reconstruction itself. (ncbi.nlm.nih.gov)
Phylogenetic trees and their interpretation
A phylogenetic tree represents relationships through branches and nodes. Tips correspond to sampled entities, such as species, individual organisms, or gene sequences. Internal nodes represent inferred ancestral lineages or divergence events at the level modeled. The branching pattern is the tree’s topology. A rooted tree identifies an ancestral direction; an unrooted tree represents relationships without specifying where the common ancestral root lies. (ncbi.nlm.nih.gov)
A clade, or monophyletic group, contains an ancestor and all its descendants. Sister groups share an immediate common ancestral node. Rotating branches around a node does not change these relationships: the order of tips across a page is a graphical choice, not evidence of relatedness. Likewise, trees do not arrange living organisms on a ladder from “primitive” to “advanced,” and a living tip is not automatically the ancestor of another living tip. (evolution.berkeley.edu)
Branch lengths require explicit interpretation:
- A cladogram emphasizes branching order; its drawn branch lengths need not carry quantitative information.
- A phylogram scales branches to estimated evolutionary change, commonly substitutions per sequence position.
- A chronogram, or time tree, scales branches to elapsed time.
These representations convey different information. A long branch in a phylogram may reflect rapid sequence evolution, a long interval of time, or both. (ncbi.nlm.nih.gov)
Root placement can use an outgroup: a lineage outside the focal group whose relationship supplies a reference for rooting. Clock assumptions and non-reversible substitution models provide other approaches. Under standard stationary, time-reversible sequence models, sequence likelihood alone generally does not distinguish alternative root positions on the same unrooted tree. (iqtree.org)
Evidence and preparation of data
Morphological and molecular characters
Morphological analyses use characters of anatomy, development, or other observable traits. They are particularly important for fossil organisms that lack recoverable molecular sequences. Molecular analyses compare DNA, RNA, or protein sequences, while genome-scale studies can also examine gene content and gene order. The expansion of DNA sequencing has greatly increased the quantity of evidence available for phylogenetic reconstruction. (pubmed.ncbi.nlm.nih.gov)
Sequence-based inference usually begins with a multiple sequence alignment, which proposes which positions are homologous across sequences. Incorrect alignment can create misleading character correspondences. Choosing appropriate genes is equally important: orthologous genes diverged through lineage splitting, whereas paralogous genes arose through duplication. Comparisons that unknowingly mix different duplicated copies can reconstruct gene-family history rather than the intended organismal relationships. (ncbi.nlm.nih.gov)
Models of evolutionary change
Molecular inference commonly uses substitution models that describe probabilities of changes between nucleotide or amino-acid states. Models may allow unequal state frequencies, different substitution rates, and variation in evolutionary rate among sites. They account probabilistically for changes that are not visible as simple present-day differences, including repeated substitutions at the same position. (ncbi.nlm.nih.gov)
A continuous-time Markov chain often describes this process. Different genes or codon positions may be assigned separate model partitions. Mixture models instead allow sites to draw from multiple evolutionary processes without requiring a fixed assignment of every site in advance. These approaches accommodate some forms of heterogeneity, but remain approximations to biological evolution. (beast.community)
Principal inference methods
Distance methods
Distance methods summarize comparisons between pairs of sequences or organisms as an evolutionary-distance matrix. Procedures such as neighbor joining then construct a tree from those distances. They can be computationally efficient, but compression into pairwise distances discards some information about the original distribution of character states. Their results also depend on how distances are estimated and whether those estimates adequately account for repeated substitutions and rate differences. (ncbi.nlm.nih.gov)
Maximum parsimony
Maximum parsimony seeks trees requiring the smallest total number of character-state changes, under specified rules for counting or weighting changes. It can be applied to morphological and molecular data. However, minimizing change is not always equivalent to recovering the correct history. Under some combinations of evolutionary rates and branch lengths, parsimony is statistically inconsistent: increasing the amount of data can strengthen support for an incorrect tree. (gen-files.princeton.edu)
Maximum likelihood
Maximum likelihood evaluates how probable the observed data would be under a proposed tree and evolutionary model. It searches for the topology, branch lengths, and model parameters that maximize the likelihood function. Probabilities are summed over unobserved ancestral character states rather than requiring a single fixed reconstruction of every ancestor. (pubmed.ncbi.nlm.nih.gov)
The number of possible trees grows rapidly with the number of sampled entities, so practical analyses generally use heuristic searches rather than exhaustive enumeration. Efficient calculations, model selection, and search strategies make likelihood inference feasible for large datasets, but finding the global optimum is not guaranteed in every analysis. (iqtree.org)
Bayesian inference
Bayesian inference combines a likelihood model with prior distributions over trees and parameters to obtain a posterior distribution. Rather than producing only one optimum, it characterizes a distribution of plausible histories conditional on the data and model. Markov chain Monte Carlo is widely used to sample this distribution. Numerical reliability depends on adequate exploration of tree space, convergence, and sufficient effective sampling. (nbisweden.github.io)
Uncertainty and statistical support
An estimated tree should be distinguished from the strength of evidence for its individual branches. In a conventional phylogenetic bootstrap, alignment columns or other character units are resampled with replacement, and trees are inferred from the resulting replicate datasets. The proportion of replicates recovering a branch measures its stability under that resampling procedure; it is not automatically the probability that the branch is historically correct. (pubmed.ncbi.nlm.nih.gov)
Bayesian branch support is typically the posterior frequency of a clade. This has a different interpretation from bootstrap support and is conditional on the chosen likelihood model and priors. Neither measure eliminates systematic error: a dataset analyzed under an inadequate model can strongly support the wrong history. Unresolved branching, often represented by a polytomy, can reflect insufficient information rather than a demonstrated simultaneous divergence. (nbisweden.github.io)
Gene trees, species trees, and networks
A gene tree describes the ancestry of a particular genetic locus or gene family; a species tree describes the divergence history of species or populations. They need not coincide, even when both are accurately inferred. Gene duplication and loss, recombination, horizontal gene transfer, and hybridization can generate differing histories across a genome. (arxiv.org)
Another source of disagreement is incomplete lineage sorting. Genetic variants present in an ancestral population can persist through successive species divergences. Their later genealogical relationships may consequently differ from the order in which species separated. Models based on coalescent theory connect these gene genealogies to population-level divergence, especially when speciation events occurred close together in time. (arxiv.org)
Phylogenomics uses genome-scale evidence to study evolutionary history. Combining many loci can improve resolution, but concatenating them into one alignment does not remove genuine differences among their histories. Species-tree approaches model or summarize that discordance. Where ancestry includes substantial genetic exchange, a phylogenetic network can represent reticulation that a strictly branching tree cannot capture. (arxiv.org)
Molecular clocks and applications
A molecular clock links sequence change to time. A strict clock assumes the same substitution rate across branches; relaxed-clock models allow rates to vary. Absolute divergence dates require additional temporal information, such as fossil calibrations, independently estimated rates, or sufficiently informative sampling dates. Date estimates therefore depend on both evolutionary assumptions and calibration evidence. (beast.community)
Phylogenies also support reconstruction of ancestral traits and investigations of gene-family evolution. In comparative biology, they help account for the fact that observations from related species are not statistically independent: inherited similarities can otherwise be mistaken for independent evidence of a relationship between traits. Phylogenetic comparative methods incorporate this shared history into statistical analysis. (arxiv.org)
In epidemiology, pathogen phylogenies contribute to studies of outbreak relationships. A pathogen tree is not, however, a direct map of transmission between hosts. Unsampled infections, within-host evolution, and the transmission of multiple genetic variants can make genealogical branching differ from transmission history. Genetic proximity alone does not establish direct transmission or its direction. (pmc.ncbi.nlm.nih.gov)
Historical development and methodological limitations
Charles Darwin used branching descent to explain relationships among species in On the Origin of Species (1859). Twentieth-century phylogenetic systematics developed explicit approaches to evaluating those relationships, with Willi Hennig’s work central to the development of cladistics. Molecular data subsequently enabled statistical reconstruction from sequences; Joseph Felsenstein’s 1981 paper provided an influential treatment of maximum-likelihood inference for DNA trees. (darwin-online.org.uk)
Important limitations include errors in sequence alignment or gene identification, uneven sampling, compositional differences among lineages, and substitution saturation that obscures earlier changes. Long-branch attraction refers to the misleading grouping of rapidly evolving lineages under some data and analytical conditions. More data can reduce random uncertainty without correcting such systematic biases. Consequently, disagreement among inferred trees may reflect either real biological differences among histories or shortcomings in sampling, models, and analysis; these explanations require separate investigation. (pmc.ncbi.nlm.nih.gov)
References
- Phylogenetic reconstruction methods: an overviewpubmed.ncbi.nlm.nih.gov
- Reconstructing trees: Cladisticsevolution.berkeley.edu
- Homologies in phylogenetic analyses--concept and testspubmed.ncbi.nlm.nih.gov
- Molecular Phylogeneticsncbi.nlm.nih.gov
- Reading trees: A quick reviewevolution.berkeley.edu
- Current approaches to whole genome phylogenetic analysispubmed.ncbi.nlm.nih.gov
- Comparative Genomics and New Evolutionary Biologyncbi.nlm.nih.gov
- Evolutionary trees from DNA sequences: a maximum likelihood approachpubmed.ncbi.nlm.nih.gov
- Evolutionary Trees from DNA Sequences: A Maximum Likelihood Approachgen-files.princeton.edu