The Protein Data Bank (PDB) is an international, publicly accessible database of three-dimensional structures of biological macromolecules. Despite its name, it contains not only proteins but also nucleic acids, including DNA, and their complexes with other molecules. Its central records describe structural models, their molecular composition, and the experiments supporting them. The archive accepts both experimentally determined structures and integrative structures that combine experimental evidence with other information. (wwpdb.org)
Origins and international organization
The PDB was established in 1971 with seven structures, creating a shared repository for results from the emerging field of structural biology. Its origins were closely connected to X-ray crystallography, which had made it possible to determine the spatial organization of biological macromolecules at atomic detail. Archiving coordinate data allowed researchers to examine structures independently of the illustrations and descriptions published in scientific papers. (cdn.wwpdb.org)
The Worldwide Protein Data Bank (wwPDB) partnership was established in 2003 to coordinate the archive internationally. Regional organizations provide deposition, curation, and access services under shared arrangements rather than maintaining unrelated collections. RCSB PDB, Protein Data Bank in Europe, and Protein Data Bank Japan are major access and deposition centers; processing responsibilities also include a center in China and specialized resources for magnetic-resonance data. (wwpdb.org)
What a structure record contains
A conventional atomic-resolution entry contains coordinates specifying the positions of individual atoms, together with identifiers for molecular entities, chains, and residues. Protein residues correspond to amino acids, while nucleic-acid residues correspond to nucleotides. Coordinate records can also describe bound ligands, ions, and water molecules. Additional fields record occupancy and displacement parameters, which help characterize how atoms are represented in the model. (pdb101-east.rcsb.org)
An entry also includes molecular names, sample sequences, source organisms, experimental methods, structure-determination details, and author information. Consequently, a PDB record represents a particular structural investigation rather than a definitive structure for an entire protein species. Different entries may describe the same protein with altered sequences, different molecular partners, or different portions of its chain. (wwpdb.org)
The deposited coordinates do not necessarily include every atom in the experimental sample. Flexible loops and chain termini may lack sufficiently clear experimental evidence to support coordinates; hydrogen atoms are also commonly absent from crystallographic and electron-microscopy models. Sequence records and coordinate records therefore serve different purposes and may have different coverage. (pdb101.west.k8s.rcsb.org)
Experimental methods and deposition
Major sources of PDB structures are crystallography, nuclear magnetic resonance spectroscopy, and cryo-electron microscopy. Each method produces different observations and requires different supporting data. Crystallographic submissions normally include structure factors or diffraction intensities; magnetic-resonance submissions include restraints and chemical shifts; three-dimensional electron-microscopy models require deposition of the corresponding map in the Electron Microscopy Data Bank (EMDB). (wwpdb.org)
The wwPDB’s OneDep system coordinates deposition, annotation, and validation across participating centers. It supports linked submission of coordinate models and experimental records, including magnetic-resonance data associated with the Biological Magnetic Resonance Data Bank (BMRB). Shared procedures check such features as ligand chemistry and consistency between the declared sample sequence and the structural model. (wwpdb.org)
Integrative structures use several kinds of evidence, potentially including traditional structural experiments, scattering, chemical crosslinking, and computational models. The archive accepts these through PDB-IHM when they are at least partly based on experimental data. Their representations can include ensembles, multiple states, or components described at different spatial scales rather than exclusively atom-by-atom models. (wwpdb.org)
Data formats and molecular assemblies
The principal archival format is PDBx/mmCIF, a structured text format supported by a data dictionary. It defines relationships among coordinates, sequences, chemical components, and experimental metadata. Unlike the older fixed-column PDB format, it does not impose the same limits on numbers of atoms, residues, or chains. The legacy format remains available for some entries but is no longer extended to accommodate new content. (pdb101-east.rcsb.org)
An important distinction is that between the crystallographic asymmetric unit and the biological assembly. The asymmetric unit supplies the independent structural information needed to generate a crystal through symmetry operations. The biological assembly represents the molecular complex shown or believed to function biologically. It may comprise all, part, or multiple copies of the asymmetric unit. For example, functional hemoglobin contains four protein chains. Assembly metadata specifies the selections and transformations needed to construct these complexes. (pdb101.rcsb.org)
Validation and interpretation
Validation reports describe model geometry and agreement with experimental observations. They can identify unusual bond geometry, close interatomic contacts, and ligand-modeling problems. For crystallographic structures, they also assess fit to electron density. These checks document potential problems at both overall and local levels; they do not make every deposited coordinate equally certain. (west.k8s.wwpdb.org)
Crystallographic resolution indicates the detail supported by the diffraction data and is commonly expressed in ångströms. Smaller numerical values generally correspond to finer detail. Resolution nevertheless describes the dataset rather than guaranteeing identical reliability throughout a model, whose flexible regions may be less clearly defined. (pdb101.rcsb.org)
Access and computational reuse
Access portals support searches by molecular names, sequences, ligands, authors, and identifiers. Programmatic interfaces expose structured metadata for computational analysis. The RCSB data service organizes information into entries, molecular entities, chain instances, assemblies, and chemical components, reflecting the archive’s underlying data relationships. (rcsb.org)
PDB structures also support computational approaches to protein folding, including AlphaFold. Access portals may display externally supplied computed structure models alongside PDB records. These predictions retain their own provenance and identifiers: appearance in the same search interface does not make a prediction an experimentally determined PDB structure. (pdb101-west.rcsb.org)