Proteins are a major class of biological macromolecules. They carry out many of the catalytic, structural, transport, signaling, and regulatory functions required by cells. Much of this functional diversity begins with the enormous number of possible amino-acid sequences. A sequence constrains how a polypeptide folds, moves, and interacts with other molecules, although its environment and binding partners can also affect the structures it adopts.
Reflecting their diverse functions, proteins also have diverse shapes and sizes. Some proteins, such as hemoglobin, are globular, whereas others, such as collagen, are fibrous. A protein’s structure is closely connected to what it can do
How different can the overall shapes of proteins be?

Different proteins can adopt very different overall shapes (source
The compact, globular examples and the elongated, fibrous examples make shape itself a biological variable rather than a cosmetic feature of the drawing.
Protein structure prediction asks a deceptively simple question:
Given a protein’s amino-acid sequence, what three-dimensional structure is it likely to have?
From a computer-science point of view, we want to map a 1D sequence to a 3D geometric object.
Background on amino acids
So what is an amino acid? An amino acid is a small organic molecule that can serve as a building block of proteins. The standard amino acids share an alpha carbon bonded to:
- An amino group (),
- A carboxyl group (),
- A hydrogen atom (), and
- A variable side chain ().
A common shorthand for the uncharged form is , although proline’s side chain closes back onto the amino nitrogen. Protonation states depend on the chemical environment; near physiological pH, free amino acids are commonly represented as zwitterions with positively charged amino () and negatively charged carboxylate groups ()
Which part of this molecular scaffold is shared, and which part is allowed to vary?

General structure of an amino acid
The side chain is what mostly distinguishes one amino acid from another. Different side chains have different sizes, charges, polarities, and chemical behaviors.
For our purposes, the important idea is simple:
A protein sequence is not a sequence of interchangeable symbols. Each symbol corresponds to a chemically different residue.
What can change at the side chain while the common amino-acid scaffold remains recognizable?

Different side chains produce different amino acids (source
Now ask the same question in three dimensions: when you replace methionine with another residue, which atoms remain part of the common backbone and which atoms change with the side chain? Use the residue codes to compare the structures, and toggle the hydrogens when they obscure the heavier-atom scaffold.
Amino-acid component
Methionine (M)
Across the choices, the amino group, , and carboxyl group remain recognizable while the side chain changes in size, branching, and chemical composition. The viewer supports geometric inspection; it does not show which conformation is most stable in a particular environment.
The peptide bond
Amino acids are connected into chains by peptide bonds. A peptide bond is the covalent bond between the carbonyl carbon of one amino acid and the amino nitrogen of the next.
For formal atom accounting, peptide-bond formation can be pictured as removing from the carboxyl group of one amino acid and from the amino group of another. The remaining atoms form a linkage and .
Which atoms are removed from the two amino acids, and which new linkage remains between the residues?

The schematic is formal atom bookkeeping: and account for the water molecule, while the carbonyl carbon and amino nitrogen remain connected in the peptide group.
Does the same atom accounting remain visible in three dimensions? Toggle between the free components and formal products, and track the carbonyl carbon, amino nitrogen, and separated water molecule.
Peptide bond
A + G
Separate wwPDB CCD component geometries for A and G. Toggle “Formal products” to compare atom accounting and peptide connectivity. Drag to rotate and scroll to zoom.
The product view adds the peptide linkage and displays separately. This comparison represents the formal reactants and products; it is not a time-resolved reaction mechanism.
Suppose we have the chain
Once amino acids have been incorporated into a chain, the units are called amino-acid residues or residues for short. In this example, methionine is one residue in the chain .
Important
A protein is not merely a string of abstract tokens. It is a chemically connected chain with bond lengths, bond angles, steric constraints, charges, and rotatable bonds. A structure-prediction model must produce a geometry that respects these coupled constraints, which is one reason the problem is difficult.
N-terminus and C-terminus
Once many residues have been linked together, the result is a polypeptide chain.
Definition.A polypeptide is a chain of amino-acid residues connected by peptide bonds. The ordered sequence of those residues is its primary structure.
A linear polypeptide chain has direction:
- The N-terminus is the end containing the terminal amino group.
- The C-terminus is the end containing the terminal carboxyl group.
- Protein sequences are conventionally written from the N-terminus to the C-terminus.
For our example,
- is the N-terminal residue.
- is the C-terminal residue.
- is residue 1, is residue 2, and so forth.
This N-terminus to C-terminus ordering is important because AlphaFold2 receives an ordered sequence, not an unordered collection of residues.
The four levels of protein structure
Before naming the structural levels, what does a folded single chain look like when we can rotate it as one three-dimensional object? The figure below shows crambin (PDB 1CRN)The Protein Data Bank, or PDB, stores experimentally determined three-dimensional structures of proteins, nucleic acids, and other biological macromolecules., a compact 46-residue plant proteinThe structure contains local patterns such as -helices and -strands, together with the global arrangement that packs them into one fold..
Crambin (PDB 1CRN)
Experimental coordinates · drag to inspect
Rotating crambin reveals local helices, strands, and loops packed into one compact global fold. This is an experimental coordinate model, so it illustrates the geometric object that structure prediction targets rather than the dynamics by which a protein folds.
The terms primary, secondary, tertiary, and quaternary describe different levels of protein organization:
- The ordered amino-acid sequence is the primary structure.
- Recurring backbone patterns, such as -helices and -strands, form secondary structure.
- The complete arrangement of one chain in 3D space is its tertiary structure.
- When multiple polypeptide chains assemble, their arrangement is the quaternary structure of the complex
[Branden et al., 1999] .
How do these four labels change the scale at which we describe the same molecular organization?

The four levels of protein structure (source
The diagram nests an ordered sequence inside local motifs, those motifs inside one chain’s fold, and folded chains inside a multi-chain assembly. These are levels of description, not four independent kinds of molecule.
Primary structure
The primary structure is the ordered 1D amino-acid sequence. Reusing our example,
This is the primary structure of the example chainYou can track the current structural level in the figure below.
.
Order matters. and are different sequences and generally correspond to different polypeptides.
What does a primary-structure diagram specify, and what geometric information does it leave unknown?

An example of protein primary structure (source
From a machine-learning point of view, the target sequence is the starting input. It provides each residue’s amino-acid identity and order. Once the residues are indexed, we can compute sequence separation, such as , but the corresponding separation in 3D is still unknown.
Secondary structure: local backbone patterns
Secondary structure refers to recurring local conformations of the protein backbone, commonly characterized by their backbone geometry and hydrogen-bonding patterns
.
Which backbone patterns recur locally, before we consider the complete fold?

Examples of protein secondary structure (source
The two most common regular patterns are:
- -helix: Notice the dotted lines in the top right of the diagram. In a standard -helix, the carbonyl oxygen of residue hydrogen-bonds with the backbone group of residue . Repeating this pattern coils a nearby stretch of the chain into a helix.
- -sheet: It looks like folded paper or a zig-zagging plane. It is formed by multiple segments of the protein chain (called -strands) lining up next to each other. Those strands can be far apart in the sequence even though they are neighbors in 3D.
The important distinction: an -helix is built from one continuous stretch of sequence, whereas neighboring -strands may come from distant parts of the chain.
Secondary-structure labels describe recurring backbone geometry and hydrogen-bonding patterns. A residue can therefore have a local role, such as belonging to a helix, and a global role determined by how that helix is positioned relative to the rest of the protein.
What do an -helix and a -sheet look like at closer range, and what holds their repeated shapes together?


The first panel contrasts the continuous winding of an -helix with neighboring -strands. The second makes the repeated backbone hydrogen bonds between strands explicit.
Tertiary structure: the complete fold of one chain
Tertiary structure is the overall three-dimensional arrangement of a single polypeptide chain. It describes how helices, sheets, loops, and side chains are positioned relative to one another
.
How are local helices, sheets, and loops packed into the global fold of one chain?

A tertiary structure is formed by arranging multiple secondary-structure elements in 3D (source
Note
If primary structure is the 1D sequence, tertiary structure is the global 3D
architecture of one chain.
Describing tertiary structure requires long-range relationships. For example:
CODE
Residue 15 ─...─...─...─...─ Residue 120These residues are far apart in sequence, but folding may bring them close together in space:
CODE
sequence distance: largespatial distance: smallThis is one of the central difficulties of protein-structure prediction. Local sequence neighborhoods are not enough. The model must infer:
- Which distant residues become spatial neighbors, and
- How they are oriented relative to one another.
That observation will later motivate one of AlphaFold2’s central design choices: representing relationships between pairs of residues, not only features attached to individual residues.
Quaternary structure: multiple chains together
Many functional proteins contain more than one polypeptide chain. Each chain has its own primary, secondary, and tertiary structure. The arrangement of multiple chains is called the quaternary structure
.
For example:
CODE
Chain A + Chain B + Chain C → protein complexIn structure files, chains are commonly labelled A, B, C, and so on. A chain is also often called a subunit when it is part of a multi-chain complex.
The original AlphaFold2 system described in 2021 primarily targeted single-chain prediction
- which chains interact
- where their interfaces lie
- how the chains are arranged
From a chain to three-dimensional geometry
The four structural levels tell us what scale we are describing. To understand the output of a structure-prediction model, we also need a more mechanical description of how the atoms are arranged.
A structure model records atomic coordinates, but those coordinates are one representation of the molecule rather than a preferred pose in space. We will approach that representation in two steps: first through the internal rotations of the chain, then through three-dimensional coordinates and local frames.
Each residue contributes three backbone atoms
- The peptide nitrogen ,
- The alpha carbon , and
- The carbonyl carbon
For residue , we will write them as , , and . The side chain branches from , while the peptide bond connects to .
Why is a useful geometric anchor for a residue?

The atom is bonded to the backbone nitrogen, the carbonyl carbon, a hydrogen, and the side chain. Except in glycine, these are four different groups, so their spatial arrangement has a definite handedness.
The three backbone neighbors locate within the chain, while the fourth substituent gives most residues a definite handed arrangement. We will later use the ordered backbone atoms to construct a residue-local coordinate frame.
A backbone with rotatable joints
Bond lengths and bond angles fluctuate, but they stay close to strongly preferred values. If we temporarily hold them fixed, most of the backbone’s remaining flexibility comes from rotations around bonds. These rotations are described by torsion angles, also called dihedral angles.
A dihedral angle is determined by four consecutively bonded atoms. The first three atoms define one plane, the last three define another, and the signed angle between those planes measures the twist around the bond shared by the middle two atoms. We will use AlphaFold’s per-residue indexing, in which the peptide-bond torsion preceding residue is called . For an internal residue , the three backbone torsions are:
- (pre-omega) uses and measures rotation around the peptide bond The IUPAC convention attaches a peptide-bond torsion to the preceding residue: uses . Therefore AlphaFold’s is the same geometric angle as IUPAC’s .
[IUPAC-IUB Commission on Biochemical Nomenclature, 1970] . - (phi) uses and measures rotation around .
- (psi) uses and measures rotation around .
With AlphaFold’s indexing:
- and are undefined because there is no preceding residue.
- is undefined because there is no following residue.
Which four atoms define each torsion angle, and which middle bond acts as its rotation axis?

Torsion angles in a protein backbone (source
The angles and provide most of the backbone’s conformational freedom, so we will focus on themThe peptide bond has partial double-bond character, which makes it nearly planar and strongly restricts . Most peptide bonds are near in the trans configuration. The cis configuration, near , is uncommon but occurs more often before proline than before other residues.
.
The backbone is not the whole structure. Most side chains have additional rotatable bonds described by side-chain torsion angles , although which angles exist depends on the residue type. Glycine and alanine have no side-chain angles, while longer side chains have one or more. Changing a angle repositions side-chain atoms without changing the backbone’s and angles
Intuition
In an idealized model, imagine the backbone as nearly rigid peptide units connected by rotatable joints. The preferred bond lengths and angles preserve local geometry, while , , and occasionally determine how neighboring units turn. Side-chain angles play the same role for atoms branching away from the backbone.
Rotating or changes the direction of the backbone, but these angles cannot vary freely. Some choices bring non-bonded atoms too close together, creating a steric clashAtoms have a finite effective size. When two non-bonded atoms are forced much closer than their preferred separation, repulsion rises steeply; structural models call such an implausibly close contact a steric clash..
The same structure can have different coordinates
At the scientific level, protein structure prediction asks for a plausible three-dimensional structure of a polypeptide chain with amino-acid sequence
where is the alphabet of standard residue types. For example, the six-residue sequence has .
What would it mean for a model to return a structure? For each residue , let be the set of heavy atoms present in residue type . The main geometric output is
In plain language, we predict a three-dimensional coordinate for every non-hydrogen atom represented in the polypeptide chain. A simplified coordinate table looks like this:
CODE
Atom x y zResidue 1 N ...Residue 1 C-alpha ...Residue 1 C ...Residue 1 O ...Residue 1 C-beta ...Other side-chain atoms ...Residue 2 N ......These numbers are measured in a global coordinate system:
- a chosen origin, together with
- three perpendicular axes.
We could rotate a structure on the screen or move it to the other side of the scene without changing any bond, angle, or spatial relationship within the molecule.
Suppose every atomic coordinate is transformed in the same way:
where is a rotation and is a translationA matrix in has orthonormal columns and determinant equal to . It preserves lengths and angles without reflecting space.. The individual coordinate triples change, but the protein’s internal geometry does not.
For example, the distance between atoms and remains:
Bond lengths, bond angles, and torsion angles are likewise unchanged. The two coordinate sets therefore describe the same structure in different global poses.
Intuition
The coordinate table depends on where we place the protein and which way it faces. Its internal geometry does not. A structure-prediction model should therefore treat two coordinate sets related by a shared rotation and translation as the same answer.
Why a reflection is not another pose
Note
Why did we require ? A matrix with orthonormal columns and determinant reverses orientation: it describes a reflection, possibly followed by a rotation. A reflection still preserves distances, so distance alone cannot tell a structure from its mirror image. It does, however, reverse handedness
An object is chiral if it cannot be placed on top of its mirror image using only rotations and translations
.
The same distinction appears in amino acids. Except for glycine, the atom in a standard amino acid is bonded to four different groups, giving the residue a particular three-dimensional handedness. Reflecting all of its coordinates reverses that arrangement. The result is the mirror-image configuration, not the original residue viewed from another direction.
Important
Rotations and translations move a structure without changing its handedness; a reflection produces mirror-image geometry. This is why two protein structures are considered equivalent under rotations and translations, not under arbitrary distance-preserving transformations.
A reflection preserves all pairwise distances. Distances alone therefore cannot distinguish a protein from its mirror image; a complete geometric description must retain some information about direction or handedness.
Let each residue carry its own coordinate system
Global coordinates contain the geometry we need, but they also contain an arbitrary choice we do not care about: a global origin and its axes. We would like to describe where an atom lies relative to a residue, in a way that does not change when the whole molecule is moved.
The three ordered, non-collinear backbone atoms , , and provide enough information to attach a local coordinate frame to residue . We can place its origin at , point one axis toward , and use to determine the backbone plane.
There are two possible directions perpendicular to that plane, so we choose one handedness convention and use it for every residue. The main AlphaFold2 article will develop the exact construction; here, we only need the idea that every residue receives its own origin and axes.
Let the frame of residue be represented by a rotation and translation . Here:
- is the global position of the frame’s origin.
- The columns of are unit vectors pointing along its three local axes.
The frame maps a point with local coordinates into global coordinates
Because is a rotation, . We can therefore express a global point in residue ‘s frame by reversing the transformation:
This equation answers a concrete geometric question:
Where is atom when residue is used as the origin and its backbone determines the axes?
Now move the whole protein by a rotation and translation . The atom and residue frame move together:
The coordinates seen from residue do not change:
The global coordinates changed, but the atom’s position relative to the residue did not. A local frame removes the arbitrary global pose while preserving directional information that distances alone would discard.
Torsion angles and local frames now have separate roles:
- Torsion angles describe how bonded parts of the chain turn around particular bonds.
- A residue-local frame describes the position and orientation of other atoms or residues from that residue’s point of view.
Can one idealized backbone connect these two ideas: how internal torsions change its shape, and why coordinates measured in a residue-local frame ignore the molecule’s global pose?
Backbone geometry
Idealized chain · shared pre-ω, φ, ψ
In the Torsion angles view, the -helix and -strand presets change and , so the highlighted four-atom definition and the resulting backbone bend change together. In Local frame, toggling the global pose rotates and translates the whole chain while the coordinates measured from the selected residue remain fixed. The figure demonstrates these geometric relationships for an idealized short backbone; it is not a simulation of protein-folding dynamics.
What a structure-prediction model must produce
A useful prediction is more than a table of coordinates. Its atoms must form the specified sequence, respect plausible local chemistry, bring the correct distant residues together, and preserve molecular handedness. Yet its absolute position and orientation do not matter: applying one rotation and translation to the entire structure gives an equivalent answer.
AlphaFold2 encodes these requirements using residue features, pair features, local frames, and torsion angles. The next article traces how those objects become parts of a concrete neural architecture.
References
- [Odden, 2020]Amino acids Backbone: Amino Acid Structure — What it is[HTML]Odden, Joanne, 2020. YouTube.
- [Reddy, 2026]Amino acid[HTML]Reddy, Michael K., 2026. Encyclopedia Britannica.
- [Mattaini, 2020]Amino Acids and Proteins[HTML]Mattaini, Katherine, 2020. Introduction to Molecular and Cell Biology (RWU BIO103).
- [Branden et al., 1999]Introduction to Protein StructureBranden, Carl-Ivar and Tooze, John, 1999. Garland Science.
- [Dunbrack Lab Team]Protein Side-Chain Conformational Analysis[HTML]Dunbrack Lab Team. Fox Chase Cancer Center.
- [Lovell et al., 2003]Structure Validation by Cα Geometry: φ,ψ and Cβ Deviation[DOI]Lovell, Simon C., Davis, Ian W., Arendall, W. Bryan, et al., 2003. Proteins: Structure, Function, and Genetics, vol. 50, pp. 437--450
- [IUPAC-IUB Commission on Biochemical Nomenclature, 1970]Abbreviations and Symbols for the Description of the Conformation of Polypeptide Chains: Tentative Rules (1969)[HTML][DOI]IUPAC-IUB Commission on Biochemical Nomenclature, 1970. Biochemistry, vol. 9, pp. 3471--3479
- [International Union of Pure et al.]International Union of Pure and Applied Chemistry. Compendium of Chemical Terminology (the Gold Book).
- [IUPAC-IUB Joint Commission on Biochemical Nomenclature, 1984]IUPAC-IUB Joint Commission on Biochemical Nomenclature, 1984. Pure and Applied Chemistry, vol. 56, pp. 595--624
- [Jumper et al., 2021]Highly Accurate Protein Structure Prediction with AlphaFold[DOI]Jumper, John, Evans, Richard, Pritzel, Alexander, et al., 2021. Nature, vol. 596, pp. 583--589
- [Biostructure, 2025]Levels of Protein Structure[HTML]Creative Biostructure, 2025. Creative Biostructure.
- [Fowler et al., 2013]Concepts of Biology[HTML]Fowler, Samantha, Roush, Rebecca, and Wise, James, 2013. OpenStax, Rice University.
- [Mao, 2021]Geometry for Computing Dihedral Angles[HTML]Mao, Lei, 2021. Lei Mao's Log Book.
