Macromolecular X-ray Crystallography
Principles of Diffraction · Crystal Growth · The Physics of Scattering
1. Introduction to Macromolecular Crystallography
Macromolecular X-ray crystallography is an advanced biophysical method used to determine the three-dimensional, atomic-resolution structures of biological macromolecules, including proteins, nucleic acids, and their complex assemblies.
1.1 The Purpose of X-ray Crystallography
An atomic-level understanding of biological systems requires resolving the positions of individual atoms in three-dimensional space. The average covalent bond length in biological molecules is approximately 0.12 nm (1.2 Å). To resolve two adjacent atoms as distinct entities, the wavelength of the radiation used for imaging must be of the same order of magnitude as the bond length — typically λ < 0.24 nm. This physical requirement places macromolecular imaging strictly in the X-ray region of the electromagnetic spectrum, with wavelengths typically ranging from 0.1 to 100 Å (with 1.54 Å from copper targets being standard for laboratory sources).
Instead, structure determination proceeds as a three-step indirect method:
Figure: The three-step crystallographic method. Step 1 — irradiate the crystal with a narrow, monochromatic X-ray beam. Step 2 — record the position and intensity of every diffracted spot (reflection) on a detector. Step 3 — computationally reconstruct the electron density, and hence the atomic model, using a Fourier transform as a "mathematical lens."
1.2 The Prerequisite of Crystallization
A single protein molecule in solution scatters X-rays uniformly in all directions, but because a single molecule contains relatively few electrons, the intensity of this scattered signal is far too weak to be detected above background noise. To overcome this limitation, the macromolecule must be organized into a crystal — billions of identical molecules arranged in a highly ordered, repeating three-dimensional array.
Figure: From molecule to lattice. A single molecule in solution scatters too weakly to detect. Organizing billions of identical molecules into an ordered, repeating lattice of unit cells allows their scattered waves to interfere constructively in specific directions, amplifying the signal by orders of magnitude into measurable diffraction spots.
2. The Art and Science of Protein Crystallization
Because crystallization is an absolute prerequisite for X-ray diffraction, growing high-quality, single crystals is the first — and often most challenging — step in structural biology. This process is frequently characterized as "more art than science," requiring extensive screening of chemical and environmental parameters.
2.1 Biophysics of Nucleation and Crystal Growth
Protein crystallization is a thermodynamic process in which a soluble protein is slowly and gently driven out of its aqueous solution to form a highly ordered solid phase, rather than a disordered, useless amorphous precipitate. This is conceptualized using a phase diagram plotting protein concentration against precipitant concentration.
Figure: Protein crystallization phase diagram. As precipitant and protein concentration rise together (e.g. via vapor diffusion), the system moves from the under-saturated zone into the metastable zone. A brief excursion into the nucleation zone spontaneously forms nuclei, which locally deplete surrounding protein concentration, dropping the system back into the metastable zone where the nucleus grows cleanly into a single large crystal. Overshooting into the precipitation zone instead yields useless amorphous aggregate.
- Under-saturated zone: the protein remains completely soluble at all concentrations here — no crystals can form.
- Metastable zone: the solution is slightly supersaturated, but the thermodynamic energy barrier is too high for spontaneous nucleus formation. If a pre-existing crystal seed is introduced here, it grows stably and cleanly.
- Nucleation zone: the solution is sufficiently supersaturated that protein molecules spontaneously overcome the thermodynamic barrier to form stable, ordered oligomeric clusters called nuclei. Nucleus formation lowers the local protein concentration, bringing the surrounding solution back down into the metastable zone, where it grows into a single large crystal.
- Precipitation zone: the solution is highly supersaturated, causing rapid, uncontrolled, disordered aggregation — useless amorphous white precipitate (flocculation) rather than structured crystals.
2.2 Precipitant Chemistry and Action
To drive a protein into the nucleation and metastable zones, precipitants are added to lower the protein's solubility in a controlled manner.
| Precipitant class | Examples | Mechanism of action |
|---|---|---|
| Ionic compounds (salts) | Ammonium sulfate [(NH4)2SO4], sodium chloride (NaCl) | "Salting out" — salt ions compete with the protein for water molecules, becoming highly hydrated and stripping the protective hydration shell from the protein surface. This exposes hydrophobic patches, forcing proteins to interact and self-assemble. |
| Organic solvents | Ethanol, isopropanol, 2-methyl-2,4-pentanediol (MPD) | Reduce the solvent's dielectric constant, increasing electrostatic attraction between oppositely charged residues on neighboring proteins. Often interact directly with the hydrophobic core, risking denaturation and unfolding. |
| Water-soluble polymers (PEG) | PEG 400, PEG 4000, PEG 8000 | The most widely used and successful precipitant class. Acts via volume exclusion (molecular crowding) — large PEG chains occupy solvent volume, physically restricting space available to the protein and forcing molecules into closer proximity without unfolding them. A weak denaturant. |
2.3 Crystallization Methodologies: Vapor Diffusion
The most common technique for achieving slow, controlled supersaturation is vapor diffusion, carried out in two main setups: hanging drop and sitting drop.
Figure: Hanging drop vapor diffusion. A droplet of protein + reservoir solution (1:1) hangs from an inverted cover slip sealed over a much larger reservoir volume. Because the droplet starts at half the reservoir's precipitant concentration, its water has a higher chemical potential; water vapor slowly evaporates from the droplet and condenses into the reservoir, gently and continuously concentrating both protein and precipitant in the droplet — guiding it from under-saturation through nucleation into stable, metastable-zone growth.
- Droplet setup: a small droplet (1–4 μL), a 1:1 mixture of purified protein solution and reservoir solution, is placed on a siliconized glass or plastic cover slip.
- Sealing: the cover slip is inverted and sealed over a well containing a much larger volume (500–1000 μL) of reservoir solution at high precipitant concentration.
- Chemical potential gradient: because the droplet contains only half the reservoir's precipitant concentration, the chemical potential of water in the droplet is higher than in the reservoir.
- Equilibration: water vapor slowly evaporates from the droplet and diffuses through the sealed air space, condensing into the reservoir, gently raising protein and precipitant concentration in the droplet and guiding it into the metastable zone for stable crystal growth.
2.4 Biological Variables and Optimization
Finding the precise conditions to produce high-quality, strongly diffracting crystals is a multi-dimensional screening puzzle.
| Variable | Typical range / role |
|---|---|
| Protein purity | >95% pure, homogeneous, and monodisperse (free of aggregates) |
| Protein concentration | Typically 5–25 mg/mL |
| pH and buffers | Modifies surface charge distribution, altering electrostatic packing contacts |
| Temperature | Constant incubation, commonly 4°C or 20°C, strongly affects solubility and nucleation kinetics |
| Ionic strength / additives | Trace metal ions (Mg2+, Zn2+, Ca2+) or small ligands stabilize specific conformations, facilitating crystal packing |
3. The Physics of X-rays and Scattering Mechanics
3.1 Electromagnetic Radiation and Generation of X-rays
X-rays are high-energy, short-wavelength electromagnetic waves. In a laboratory setting, they are generated using an X-ray tube or a rotating anode generator.
Figure: Generation of laboratory X-rays. A tungsten filament (cathode) is heated, emitting free electrons by thermionic emission; a large potential (30,000–50,000 V) accelerates them into a target metal (anode). Deceleration on impact produces a continuous Bremsstrahlung spectrum, while displacement of an inner (K-shell) electron — refilled by an outer (L-shell) electron — emits a sharp, monochromatic characteristic peak (Kα; 1.5418 Å for copper).
- Thermionic emission: a tungsten filament (cathode) is heated by an electrical current, emitting a stream of free electrons.
- Acceleration: a massive electric potential difference (typically 30,000–50,000 V) accelerates these electrons toward a target metal block (anode) made of copper, molybdenum, or chromium.
- Bremsstrahlung: when high-speed electrons collide with the target atoms, they decelerate rapidly, releasing kinetic energy as a broad continuous spectrum of radiation ("deceleration radiation").
- Characteristic radiation: a high-energy electron may also displace an inner-shell (K-shell) electron; an electron from a higher orbital (L-shell) immediately drops down to fill the vacancy, emitting its excess energy as a highly monochromatic characteristic X-ray photon. For a copper target, this Kα transition generates a precise wavelength of λ = 1.5418 Å (0.154 nm).
3.2 Scattering of Electromagnetic Waves by Electrons
When an X-ray wave strikes an atom, it physically interacts with the charged particles within it.
| Particle | Response to incoming X-ray field |
|---|---|
| Nucleus | Protons and neutrons are extremely massive relative to electrons, so their inertial resistance is too high to be moved by the oscillating electric field — the nucleus does not scatter X-rays. |
| Electron cloud | Lightweight electrons oscillate in resonance with the field of the incoming wave, acting as miniature antennas that re-emit (scatter) secondary electromagnetic waves of identical wavelength in all directions. |
X-rays are therefore scattered exclusively by electrons, making X-ray crystallography fundamentally a map of the electron density within the crystal. An atom's scattering power (its atomic scattering factor, f) is directly proportional to its local electron density — and thus its atomic number, Z.
3.3 Wave Interference: Constructive vs. Destructive
Because all atoms in a crystal lattice scatter X-rays, the final diffraction pattern is determined by how these scattered secondary waves combine in space.
Figure: Constructive vs. destructive interference. If two scattered waves travel paths differing by an integer multiple of the wavelength (1λ, 2λ, ...), they remain in phase — peaks align with peaks, amplitudes add, and a bright reflection is recorded. If their paths differ by a half-integer multiple (0.5λ, 1.5λ, ...), peaks align with troughs, amplitudes cancel completely, and no signal is recorded at the detector.
4. Bragg's Law and Interplanar Spacing
In 1912, W. L. Bragg simplified the complex mathematics of three-dimensional lattice scattering by proposing that diffraction can be visualized as the "reflection" of X-rays from parallel, evenly spaced planes of atoms passing through the crystal lattice.
4.1 Derivation of Bragg's Law
Consider a set of parallel lattice planes separated by a constant interplanar distance d. A monochromatic beam of X-rays with wavelength λ strikes these planes at an incident angle θ (the Bragg angle).
Figure: Bragg reflection geometry. Ray P strikes Plane 1 at atom O and reflects at angle θ. The parallel Ray Q strikes Plane 2 at atom A and reflects at the same angle θ, travelling an extra distance before and after reflection. Dropping perpendiculars from O onto Ray Q gives feet B and D; the right triangles OBA and ODA each have hypotenuse d, so BA = AD = d sinθ. For the two reflected waves to stay in phase, this total extra path (2d sinθ) must equal a whole number of wavelengths.
- Setup: two parallel lattice planes separated by interplanar distance d; a monochromatic beam of wavelength λ strikes both at angle θ.
- In-phase requirement: for the two reflected waves to interfere constructively, the extra distance travelled by the lower ray must equal an integer number of wavelengths (nλ).
- Geometry: dropping perpendiculars from O onto Ray Q gives right triangles with hypotenuse d, so each leg (BC and CD) equals d sinθ.
- Result: substituting the path difference (BC + CD = 2d sinθ) into the in-phase condition yields Bragg's Law.
4.2 Calculating Interplanar Spacing (d)
Bragg's Law is a fundamental tool because it links the measurable diffraction angle (θ) directly to the physical distance between atomic planes in the crystal (d). Rearranging for d (assuming n=1):
| Detector position | Diffraction angle | Interplanar spacing (d) | Structural information |
|---|---|---|---|
| Outer edge | High θ | ≈ 1.5–2.0 Å | High-resolution detail (individual atoms) |
| Near center | Low θ | ≈ 6.0–10.0 Å | Low-resolution overall shape |
4.3 Calculating Reflection Angles (2θ)
In a real diffraction experiment, the diffraction angle isn't measured directly — instead, crystallographers measure the physical distance (r) from the central, undiffracted beam to a diffracted spot on a flat detector, along with the crystal-to-detector distance (A).
Figure: Measuring the reflection angle. The angle between the primary, undiffracted beam and a diffracted beam is exactly 2θ. Measuring the spot's distance from the beam center (r) and the crystal-to-detector distance (A) allows 2θ — and hence θ and d — to be calculated for every reflection.
Dividing by 2 yields the Bragg angle θ, which is then substituted into Bragg's Law to calculate the interplanar spacing d of the atomic planes that generated that specific spot.
5. The Structure Factor and the Fourier Transform
A crystal contains billions of atoms. While Bragg's Law tells us the geometric conditions required for diffraction to occur, it does not tell us the intensity of the diffraction spots. The intensity of each spot is determined by the specific arrangement of atoms inside the unit cell.
5.1 The Unit Cell and Miller Indices
The unit cell is characterized by three edge lengths (a, b, c) and three inter-edge angles (α, β, γ). Every spot (reflection) in a diffraction pattern corresponds to a specific set of imaginary planes slicing through the unit cell, defined by three integers called Miller Indices (h, k, l).
Figure: The unit cell and a Miller plane. The unit cell's shape is fully defined by three edge lengths (a, b, c) and three angles (α between b & c, β between a & c, γ between a & b). The shaded plane represents one member of a family of parallel imaginary planes indexed by the Miller indices (h, k, l), where h, k, and l describe how many parts the plane divides edges a, b, and c into respectively.
5.2 Mathematical Form of the Structure Factor F(hkl)
Each diffracted beam corresponding to Miller indices (h, k, l) is mathematically represented as a wave called the structure factor, F(hkl) — the vector sum of all individual waves scattered by each of the N atoms inside a single unit cell.
Because F(hkl) is a complex number, it can be written in polar form:
5.3 Fourier Synthesis and Electron Density ρ(x,y,z)
The physical structure of the protein (its electron density map, ρ(x,y,z)) and the structure factors (F(hkl)) are Fourier transform pairs. If all structure factors are known (both amplitudes and phases), an inverse Fourier transform — Fourier synthesis — calculates the electron density at any coordinate in the unit cell:
6. The Phase Problem and Solution Strategies
To perform the Fourier synthesis and map the electron density, two values are needed for every reflection (h, k, l): the amplitude |F(hkl)| and the phase angle φ(hkl).
6.1 Defining the Phase Problem
Figure: The phase problem. A diffraction experiment records only intensities, from which amplitudes |F| are easily obtained. But the detector is a square-law device — it measures photon energy, not the timing of the wave oscillations — so all phase information φ is destroyed. Three major strategies exist to recover the missing phases: isomorphous replacement, anomalous dispersion, and molecular replacement.
6.2 Isomorphous Replacement (SIR and MIR)
- Method: the native protein crystal is soaked in a solution containing heavy metal atoms (mercury, platinum, gold, or uranium), which bind to specific residues (like cysteine or methionine) on the protein surface.
- Isomorphism constraint: the heavy metal must bind without altering the overall shape, packing, or unit cell dimensions of the crystal.
- Phasing mechanism: because the heavy atom has a huge atomic number, it scatters X-rays extremely strongly, significantly altering many spot intensities. Comparing native vs. derivative intensities lets crystallographers locate the heavy atoms and mathematically triangulate the protein's phase angles.
6.3 Anomalous Dispersion (SAD and MAD)
- Method: when the incident X-ray wavelength is tuned close to a heavy atom's natural absorption edge, that atom absorbs energy, causing a phase delay and a change in scattering power.
- Phasing mechanism: this anomalous absorption breaks Friedel's Law (that reflection (h,k,l) must equal the intensity of its symmetry mate (−h,−k,−l)), generating small intensity differences between "Bijvoet pairs" that are used to calculate phases.
| Variant | Approach |
|---|---|
| MAD (multi-wavelength) | Crystal irradiated with synchrotron X-rays at three or four distinct wavelengths around the target atom's absorption edge. |
| SAD (single-wavelength) | Phases calculated from precise anomalous-difference measurements at one optimized wavelength. |
| SeMet phasing | Proteins expressed with selenomethionine in place of methionine; selenium's accessible absorption edge allows direct MAD/SAD phasing without heavy-metal soaking. |
6.4 Molecular Replacement (MR)
A purely computational strategy, used when a homologous protein structure (typically >30% sequence identity) is already deposited in the Protein Data Bank.
- Search model: the known homologous structure is used as a template.
- Computational search: software rotates and translates this model within the new crystal's unit cell, calculating a theoretical diffraction pattern for each orientation.
- Alignment: when the search model's orientation matches the real protein, calculated and observed diffraction patterns align, and the resulting phases seed electron density map building for the new structure.
7. Model Building, Refinement, and Validation
7.1 Translating Electron Density into Atomic Models
Once initial phase estimates are combined with the experimental amplitudes, Fourier synthesis produces the first three-dimensional electron density map, into which the crystallographer fits the known amino acid sequence.
Figure: Model building. The known primary amino acid sequence is fitted into the contours of the calculated electron density map, tracing the peptide backbone and positioning each side chain.
| Resolution | What the map reveals |
|---|---|
| High resolution (<2.0 Å) | Sharp density — individual side chains, water molecules, even carbonyl oxygen orientations clearly resolved. |
| Medium resolution (2.5–3.5 Å) | Main polypeptide backbone resolved, but side-chain details can be ambiguous. |
7.2 Computational Refinement: R-factor and Rfree
The initial atomic model is always imperfect. Refinement is an iterative process that adjusts atomic coordinates (x, y, z) and thermal vibration parameters (B-factors) to minimize the difference between observed and calculated structure factor amplitudes, combining stereochemical restraints with the experimental X-ray data.
| R-factor value | Interpretation |
|---|---|
| ≈ 0.59 | A completely random, incorrect model. |
| 0.15–0.20 | A well-refined, correct macromolecular structure at standard resolution (80–85% match to experimental data). |
7.3 Steric Quality Control: The Ramachandran Plot
Because peptide bonds are planar, the polypeptide backbone's conformation is defined solely by two dihedral angles: φ (phi, rotation around N–Cα) and ψ (psi, rotation around Cα–C). Steric hindrance between backbone and side-chain atoms permits only specific (φ,ψ) combinations.
Figure: Ramachandran plot. Plotting ψ against φ for every residue reveals favored regions (dark) corresponding to standard secondary structures — the α-helix and β-sheet regions dominate, with a smaller left-handed helix region. Allowed regions (lighter halo) permit minor steric contacts. Residues falling outside these regions (disallowed, e.g. the flagged outlier) indicate steric clashes and must be manually re-examined.
8. Fiber Diffraction and Helical Structures
Some of the most important biological macromolecules — collagen, muscle fibers, flagella, and double-stranded DNA — are highly fibrous and elongated, and do not form ordered three-dimensional single crystals.
8.1 Principles of Non-Crystalline Diffraction
To study these materials, researchers use fiber diffraction. A sample of aligned fibers — long axes parallel to one another, but rotated randomly around those axes — is placed in the X-ray beam. When irradiated perpendicular to the fiber axis, the molecules produce distinctive, symmetric diffraction patterns that reveal their molecular dimensions.
8.2 Helical Parameters and the Famous DNA "X" Diffraction Pattern
Helical molecules like double-stranded B-DNA possess a highly specific, repeating helical symmetry. In 1953, Rosalind Franklin's famous Photo 51 demonstrated the power of fiber diffraction.
Figure: The B-DNA "X" fiber diffraction pattern. Helical diffraction produces spots arranged along diagonal lines forming a giant "X," whose angle relates to the helix's pitch angle. Reflections organize into horizontal layer lines, spaced inversely to the helical pitch (p = 34 Å for B-DNA, the distance for one full turn). Strong meridional reflections at the very top and bottom correspond to the axial rise per residue (h = 3.4 Å, the spacing between stacked base pairs).
- The "X" pattern: diffraction spots arranged along diagonal lines forming a giant X shape; the angle of the X relates directly to the helix's pitch angle (tilt).
- Layer lines: reflections organize into parallel horizontal bands. Their vertical spacing is inversely proportional to the helical pitch (p = 34 Å), the distance for one complete helical turn.
- Meridional reflections: strong reflections at the absolute top and bottom of the vertical axis; their distance from center corresponds to the axial rise per residue (h = 3.4 Å), the spacing between adjacent stacked base pairs.
- Missing reflections (systematic absences): the absence of reflections on specific layer lines reveals the presence of a double helix — two intertwined helical chains, exposing the major and minor grooves of B-DNA.
| Parameter | Symbol | Value (B-DNA) |
|---|---|---|
| Helical pitch (one full turn) | p | 34 Å |
| Axial rise per residue | h | 3.4 Å |
LessonStep 11 of 14

