Amino Acids and Peptides

1. Structure, Isomerisation, and Chemical Nature of Amino Acids

Exploring the molecular architecture, cyclic anomalies, and non-proteinogenic variations of amino acids.

The Molecular Architecture of the α-Carbon Core

Amino acids serve as the fundamental monomeric building blocks of all proteins. Each α-amino acid is organised around a central tetrahedral carbon atom, designated the α-carbon (Cα). Covalently bound to this asymmetric Cα core are four distinct substituents:

  • An acidic carboxyl group (-COOH).
  • A basic amino group (-NH2).
  • A unique, variable side chain (the R group) that dictates chemical specificity.
  • A solitary hydrogen atom (-H).

The general chemical formula of a standard α-amino acid in its non-ionised form is NH2-CHR-COOH. However, because the primary amine and carboxylic acid groups are attached directly to the same Cα, these chemical moieties mutually influence one another's ionisation states, giving rise to unique acid-base dynamics under physiological conditions.

General Structure of an α-Amino Acid

Cα COO- α-carboxyl group R Side chain (R group) H3N+ α-amino group H

Primary vs. Secondary (Imino) Amino Acids: The Structural Anomaly of Proline

Of the twenty standard, protein-encoded amino acids, nineteen possess a primary amino group (-NH2) attached to the Cα. The sole exception to this structural paradigm is proline.

Proline does not contain a free primary amino group; rather, its side chain is a rigid, five-membered pyrrolidine ring that is covalently bonded to both the Cα and the nitrogen of the amino group. Consequently, proline contains a secondary amino group (an imine, -NH-), making it technically an imino acid.

The Cyclic Architecture of Proline

Cα CH2 CH2 CH2 H2N+ COO- H ← Secondary nitrogen locked in ring

This cyclic architecture imposes profound steric constraints on the polypeptide backbone, limiting its rotational freedom and significantly influencing the local folding of protein secondary structures.

Non-α-Amino Acids: Physiological and Non-Proteinogenic Forms

While proteins are composed exclusively of α-amino acids, biologically active, non-proteinogenic amino acids exist in which the amino group is positioned on a carbon atom other than the Cα. These are designated as β-, γ-, δ-, or ε-amino acids, depending on the precise carbon atom hosting the amine moiety relative to the carboxylate terminus.

Carbon Chain Designations in Amino Acids

α-Amino Acid: H3N+ CH(R) COO- β-Amino Acid: H3N+ CH(R) CH2 COO- (α) (β) γ-Amino Acid: H3N+ CH(R) CH2 CH2 COO- (α) (β) (γ)

Examples of Physiologically Critical Non-α-Amino Acids:

  • β-Alanine: A precursor in the biosynthesis of the vitamin pantothenic acid (Vitamin B5) and a vital constituent of the dipeptide carnosine.
  • γ-Aminobutyric Acid (GABA): A major inhibitory neurotransmitter natively operating in the mammalian central nervous system, synthesised directly via the enzymatic decarboxylation of L-glutamate.

2. Optical Properties and Absolute Configurations

Examining the stereochemistry, optical activity, and formal nomenclature systems (DL and RS) governing amino acid chirality.

Optical Activity and Polarimetry

All amino acids incorporated into ribosomally synthesised proteins, with the sole exception of glycine, contain at least one asymmetric center—the α-carbon (Cα). Glycine is strictly achiral because its R group is a single hydrogen atom, meaning the Cα is bonded to two identical substituents (-H), creating a definitive plane of symmetry.

All other nineteen standard amino acids are entirely chiral molecules; they inherently lack a plane or center of symmetry and are therefore optically active. When a solution of a chiral amino acid is placed in a polarimeter, it actively rotates the plane of plane-polarised light by a characteristic angle, θ.

Polarimeter Optical Pathway Schematic

Light Source Unpolarised Light Polariser Plane-polarised Light Chiral Soln Sample Tube θ Rotated Light

Optically active molecules are strictly classified by the specific direction in which they rotate polarised light:

  • Dextrorotatory (d- or +): Rotates the plane of light to the right (clockwise).
  • Levorotatory (l- or -): Rotates the plane of light to the left (counterclockwise).

Biot's Law and Specific Rotation

The magnitude of optical rotation depends heavily on the concentration of the chiral solute, the temperature, the wavelength of light used, the solvent, and the physical path length of the sample tube. This core relationship is mathematically quantified by Biot's Law:

α = [α]Tλ × C × l

Where:

  • α = the observed optical rotation in degrees (°).
  • C = the concentration of the optically active solute in g/ml.
  • l = the physical path length of the sample cell in decimeters (dm, where 1 dm = 10 cm).
  • [α]Tλ = the specific rotation of the compound. This is a fundamental physical constant defined as the observed rotation when plane-polarised light at a specific wavelength λ (conventionally the yellow sodium D line, λ = 589 nm) passes through a sample cell of 1 dm path length containing a solute concentration of 1 g/ml at temperature T (°C).

The DL Stereochemical System

The absolute configuration of chiral amino acids was historically defined using the DL system, which correlates the spatial arrangement of the substituents around the Cα directly to the reference three-carbon aldose sugar, glyceraldehyde.

In a standard Fischer projection, the most highly oxidised carbon (the carbonyl/carboxyl group) is oriented exactly at the top of the vertical axis, and the carbon chain extends downwards:

  • If the principal functional group (the hydroxyl group in glyceraldehyde, or the amino group in amino acids) is positioned on the left side of the horizontal axis, the molecule is assigned the L-configuration.
  • If the functional group is on the right side, the molecule is assigned the D-configuration.

Fischer Projections: Glyceraldehyde vs. Alanine

L-Glyceraldehyde CHO CH2OH HO H L-Alanine COO- CH3 H3N+ H D-Glyceraldehyde CHO CH2OH H OH D-Alanine COO- CH3 H NH3+

All chiral amino acids that are ribosomally incorporated into proteins consistently exhibit the L-configuration. D-amino acids are strictly excluded from ribosomal translation, although they occur naturally in non-protein structural biopolymers, such as the peptidoglycan cell walls of Gram-positive and Gram-negative bacteria, and in short peptide antibiotics synthesised via non-ribosomal peptide synthetase (NRPS) pathways.

The Cahn-Ingold-Prelog (RS) Priority Rules

While the DL system is highly convenient for carbohydrates and amino acids, it remains an arbitrary comparative system. The modern, systematic framework for defining absolute configuration is the Cahn-Ingold-Prelog (RS) system, which firmly assigns priority to the four substituents bonded to an asymmetric carbon based on their atomic numbers.

Priority Determination Rules

  1. Atoms directly bonded to the chiral center are ranked by decreasing atomic number (-S > -O > -N > -C > -H).
  2. If two or more attached atoms are perfectly identical, priority is determined by sequentially comparing the atomic numbers of the next atoms in their respective chemical chains.
  3. Double and triple bonds are rigorously treated as if they were bonded to two or three of those individual atoms via single bonds (e.g., a carboxyl group -COO- is treated mathematically as a carbon completely bonded to three oxygen atoms).

The priority sequence for standard amino acid substituents is typically:

-SH > -OH > -NH2 > -COOH > -CH2OH > -CH3 > -H

RS System Three-Dimensional Projection (L-Alanine)

Cα (1) NH3+ (2) COO- (3) CH3 (4) H Points away Direction (1 → 2 → 3): Counterclockwise (S)

Assigning the Configuration

To accurately assign the configuration:

  1. The molecule is viewed with the lowest priority group (typically hydrogen, Priority 4) pointing directly away from the observer.
  2. The remaining three substituents (Priorities 1, 2, and 3) are visually traced in descending order:
    • If the path from 1 → 2 → 3 runs exactly clockwise, the configuration is R (rectus, right).
    • If the path runs exactly counterclockwise, the configuration is S (sinister, left).

Under this rigorous system, almost all L-amino acids normally found in proteins exhibit the S configuration at the Cα (for example, L-alanine is strictly (S)-alanine). The single, critical biochemical exception is L-cysteine. Because the heavy sulfur atom in the thiol side chain (-CH2SH) inherently possesses a higher atomic number than the oxygen atoms of the carboxyl group (-COO-), the priority rankings actively shift:

  1. -NH3+ (Priority 1)
  2. -CH2SH (Priority 2)
  3. -COO- (Priority 3)
  4. -H (Priority 4)

When traced from 1 → 2 → 3, the path visually runs clockwise, fundamentally designating L-cysteine as (R)-cysteine.

3. Taxonomy and Structural Classification of Amino Acids

Exploring the diversity, rare forms, and chemical categorization of protein building blocks.

Proteinogenic (Standard) vs. Non-Standard Amino Acids

Of the hundreds of amino acids existing in nature, only 22 are classified as standard (proteinogenic) amino acids, meaning they are specified by genetic codons and incorporated into nascent polypeptide chains during translation.

Non-standard amino acids are those that are not directly incorporated via standard ribosomal translation. Instead, they are generated through highly regulated post-translational modifications (PTMs) of standard residues already embedded within a polypeptide chain, or they function as key metabolic intermediates.

  • 4-Hydroxyproline: Generated by the post-translation modification of proline residues by prolyl 4-hydroxylase; essential for the thermal stability of the collagen triple helix.
  • 5-Hydroxylysine: Synthesised by lysyl hydroxylase; critical for the covalent cross-linking of collagen fibers.
  • Desmosine: A complex, tetra-functional amino acid formed by the condensation of four lysine side chains, providing the elastic, rubber-like properties of elastin.
  • γ-Carboxyglutamate: Synthesised via a Vitamin K-dependent carboxylation of glutamate; contains two adjacent carboxyl groups in its side chain, enabling high-affinity binding of calcium ions (Ca2+) in clotting factors like prothrombin.
  • N-Formylmethionine (fMet): The initiator amino acid in prokaryotic translation and eukaryotic mitochondrial/chloroplast organellar protein synthesis.

Non-Standard Amino Acid Structures

4-Hydroxyproline HO-CH CH2 H2C CH-COO- N+—H2 γ-Carboxyglutamate COO- COO- CH (γ) CH2 (β) H3N+ C H (α) COO-

The 21st and 22nd Standard Amino Acids

The structural taxonomy of proteinogenic amino acids is extended by two rare, specialised residues that are co-translationally incorporated in response to specific stop codons through unique transfer RNA (tRNA) reprogramming:

  • Selenocysteine (Sec, U): Known as the "21st amino acid," its structure is a homologue of cysteine in which the sulfur atom is replaced by a selenium atom (–CH2SeH). It is incorporated during translation in response to the stop codon UGA. This requires a specific downstream mRNA stem-loop structure called the SECIS (Selenocysteine Insertion Sequence) element, which recruits a dedicated translation elongation factor (SelB) and a unique tRNA (tRNASec). Selenocysteine is a catalytic residue in vital redox enzymes, including glutathione peroxidase and formate dehydrogenase.
  • Pyrrollysine (Pyl, O): The "22nd amino acid," found in methanogenic archaea and certain anaerobic bacteria. It consists of a lysine residue covalently linked to a pyrroline ring. It is incorporated at the stop codon UAG via a specialised pyrrollysyl-tRNA synthetase and a dedicated tRNAPyl, and is critical for the catalytic function of methyltransferases in methanogenesis.

Specialized Standard Amino Acids

Selenocysteine (Sec) COO- H3N+ C H (α) CH2 SeH Pyrrollysine (Pyl) COO- H3N+ C H (α) (CH2)4 NH C = O HC == N H3C-CHCH2

Structural Categorisation by R-Group Chemistry (at Physiological pH 7.4)

The physical, chemical, and structural roles of the 20 standard amino acids are determined entirely by the properties of their side chains (R groups). At physiological pH (~7.4), they are classified into four primary categories:

  1. Non-Polar, Hydrophobic Side Chains (9 Amino Acids)

    These side chains lack polar functional groups, possess low dielectric constants, and tend to cluster within the interior of folded proteins to avoid contact with the aqueous solvent, driving the hydrophobic collapse of proteins.

    • Glycine (Gly, G): R = –H. The smallest, most flexible amino acid; lacks chirality.
    • Alanine (Ala, A): R = –CH3. A simple methyl group.
    • Valine (Val, V): R = –CH(CH3)2. A branched-chain aliphatic hydrocarbon.
    • Leucine (Leu, L): R = –CH2–CH(CH3)2. Isomeric with isoleucine.
    • Isoleucine (Ile, I): R = –CH(CH3)–CH2CH3. Contains a second chiral center at the β-carbon.
    • Proline (Pro, P): R = –CH2–CH2–CH2, forming a secondary imine pyrrolidine ring.
    • Methionine (Met, M): R = –CH2–CH2–S–CH3. An ether-linked thioether.
    • Phenylalanine (Phe, F): R = –CH2–C6H5. A highly hydrophobic phenyl ring.
    • Tryptophan (Trp, W): R = –CH2–indole ring. A bulky, bicyclic aromatic system.
  2. Polar, Uncharged (Hydrophilic) Side Chains (6 Amino Acids)

    These side chains possess electronegative heteroatoms (oxygen, nitrogen, or sulfur) that can engage in hydrogen bonding with water or other polar residues, but they do not carry a net charge at neutral pH.

    • Serine (Ser, S): R = –CH2OH. Contains a primary hydroxyl group.
    • Threonine (Thr, T): R = –CH(OH)–CH3. Contains a secondary hydroxyl and a second chiral center.
    • Cysteine (Cys, C): R = –CH2SH. Contains a highly reactive thiol (sulfhydryl) group, capable of forming covalent disulfide bonds (bridges) under oxidising conditions.
    • Asparagine (Asn, N): R = –CH2–CONH2. The amide derivative of aspartate.
    • Glutamine (Gln, Q): R = –CH2–CH2–CONH2. The amide derivative of glutamate.
    • Tyrosine (Tyr, Y): R = –CH2–phenol ring. An amphipathic aromatic residue with a weakly acidic hydroxyl group.
  3. Positively Charged (Basic) Polar Side Chains (3 Amino Acids)

    These side chains contain nitrogenous bases that are protonated and carry a net positive charge at physiological pH.

    • Lysine (Lys, K): R = –CH2–CH2–CH2–CH2–NH3+. Terminated by a flexible primary ε-amino group.
    • Arginine (Arg, R): R = –CH2–CH2–CH2–NH–C(NH2)=NH2+. Terminated by a highly basic guanidinium group, which remains protonated across almost the entire physiological pH spectrum.
    • Histidine (His, H): R = –CH2–imidazole ring. Contains an weakly basic imidazole group. With a pKa close to neutral pH, it can readily transition between protonated (positive) and deprotonated (neutral) states, serving as an essential proton donor/acceptor in enzyme-catalysed reactions.
  4. Negatively Charged (Acidic) Polar Side Chains (2 Amino Acids)

    These side chains contain carboxyl groups that are fully deprotonated and carry a net negative charge at physiological pH.

    • Aspartate (Asp, D): R = –CH2–COO. A β-carboxyl group.
    • Glutamate (Glu, E): R = –CH2–CH2–COO. A γ-carboxyl group.

4. Ultraviolet Spectroscopy of Aromatic Amino Acids

Understanding the optical physics and analytical applications of conjugated π-electron systems in proteins.

The Physics of UV Absorption in Proteins

The three aromatic amino acids—phenylalanine, tyrosine, and tryptophan—contain conjugated π-electron systems that can undergo electronic transitions upon absorbing ultraviolet (UV) radiation. This property is heavily exploited for the non-destructive detection, characterisation, and quantitative measurement of proteins in biochemical solutions.

UV Absorbance Spectrum of Aromatic Amino Acids

280 nm (Assay Wavelength) Absorbance Wavelength (nm) 240 250 260 270 280 290 300 Tryptophan (Trp, λmax = 279.8 nm) Tyrosine (Tyr, λmax = 274.6 nm) Phenylalanine (Phe, λmax = 257.4 nm)

The specific spectral properties of the three aromatic residues are:

  • Phenylalanine: Exhibits a weak absorption maximum at 257.4 nm due to its single phenyl ring. Its molar extinction coefficient (ε) is exceptionally low, making it a negligible contributor to the overall UV absorbance of typical proteins.
  • Tyrosine: Features a strong absorption peak at 274.6 nm, originating from its phenolic ring. At high pH (>10), deprotonation of the phenolic hydroxyl group shifts the absorption maximum bathochromically (red-shift) to longer wavelengths and significantly increases the extinction coefficient.
  • Tryptophan: Displays the most intense absorption profile, with a major peak at 279.8 nm. This dominant absorption is directly due to its highly conjugated, bicyclic heterocyclic indole ring.

The 280 nm Spectroscopic Assay

In laboratory biochemistry, protein concentration is routinely measured by continuously monitoring absorbance at 280 nm. This specific wavelength represents a highly analytical "sweet spot" where the collective absorption of tryptophan and tyrosine dominates the spectrum:

  • Phenylalanine does not absorb significantly at 280 nm.
  • Tryptophan's absorbance at 280 nm is approximately four times greater than that of tyrosine on a molar basis, meaning the total A280 of a protein is heavily determined by its relative tryptophan content.

The total protein concentration is calculated from the measured A280 value using the fundamental Beer-Lambert Law:

A = ε × c × b

Where:

  • A = measured optical absorbance (dimensionless).
  • ε = molar extinction coefficient (M-1 cm-1), which can be calculated theoretically based on the precise number of tryptophan and tyrosine residues embedded in the sequence:
    ε280 ≈ (nTrp × 5500) + (nTyr × 1490)
  • c = molar concentration of the protein in the solution (M).
  • b = optical path length of the sample cuvette (typically strictly standardized to 1 cm).

5. Acid-Base Chemistry and Titration Dynamics

Analyzing the amphoteric properties, zwitterion formation, and quantitative titration curves of amino acids under varying physiological conditions.

Amphiprotic (Amphoteric) Nature and Zwitterions

Amino acids are classic examples of amphiprotic (or amphoteric) substances; they contain both proton-donating acidic groups (-COOH) and proton-accepting basic groups (-NH2), allowing them to act as either acids or bases depending on the pH of the surrounding solution.

At physiological pH (~7.4), the carboxylic acid group (-COOH) is fully deprotonated to its conjugate base form (-COO), while the primary amino group (-NH2) is protonated to its conjugate acid form (-NH3+). This dipolar ion species, which carries equal and opposite charges on different atoms and has a net charge of zero, is called a zwitterion (from the German Zwitter, meaning "hybrid").

Ionisation States of an Amino Acid Across the pH Scale

Low pH (Highly Acidic) COOH H3N+ C H R Net Charge: +1 Physiological pH (~7.4) COO- H3N+ – C – H R Net Charge: 0 High pH (Highly Basic) COO- H2N – C – H R Net Charge: -1

Quantitative Titration of a Non-Ionisable R-Group (e.g., Alanine)

For amino acids with simple aliphatic or uncharged side chains, there are only two ionisable groups: the α-carboxyl group and the α-amino group. If we titrate the fully protonated cationic form of alanine (at pH < 1.0) with a strong base (such as NaOH), the proton-release occurs in a stepwise, two-stage process:

0 2 4 6 8 10 12 pH 0.0 1.0 2.0 Equivalents of OH− added pK1 = 2.34 pI = 6.01 pK2 = 9.69 Ala+ Ala-

The First Dissociation Stage (Deprotonation of the α-Carboxyl Group)

At highly acidic pH, the dominant species is the fully protonated cation, Ala+ (net charge +1). As hydroxide ions (OH) are added, the highly acidic proton of the carboxyl group is released first:

Ala+ + OH ⇌ Ala0 + H2O

The midpoint of this first titration stage is reached when exactly 0.5 equivalents of base have been added. At this point, the concentrations of the cation (Ala+) and the zwitterion (Ala0) are equal ([Ala+] = [Ala0]). The pH at this midpoint is defined as the pK1 (or pKa of the carboxyl group), which for alanine is 2.34. This region exhibits strong buffering capacity, as described by the Henderson-Hasselbalch equation:

pH = pK1 + log([Ala0] / [Ala+])

The Second Dissociation Stage (Deprotonation of the α-Amino Group)

As the titration continues past pH = 6.0, the zwitterion (Ala0) is the dominant species. With further addition of base, the weaker acid (the protonated primary amino group, -NH3+) begins to lose its proton:

Ala0 + OH ⇌ Ala + H2O

The midpoint of this second stage occurs when 1.5 equivalents of base have been added. Here, the concentration of the zwitterion (Ala0) equals that of the anionic species (Ala). The pH at this midpoint is defined as pK2 (or pKa of the amino group), which for alanine is 9.69.

Calculation of the Isoelectric Point (pI)

The isoelectric point (pI) is defined as the precise pH at which the net electric charge of the amino acid population is exactly zero. At this pH, the molecule is electrophoretically immobile (it will not migrate in an electric field) and displays its minimum solubility in water, as it lacks a net repulsive charge.

For amino acids with non-ionisable side chains, the zwitterion is the intermediate species flanking the protonation states of the carboxyl and amino groups. Therefore, the pI is mathematically derived as the simple arithmetic mean of the two pKa values:

pI = (pK1 + pK2) / 2

For alanine:

pI = (2.34 + 9.69) / 2 = 6.015 ≈ 6.02

Quantitative Titration of an Acidic R-Group (e.g., Glutamate)

Amino acids with an acidic side chain contain three ionisable groups: the α-carboxyl group, the γ-carboxyl group of the side chain, and the α-amino group. The respective pKa values for glutamic acid are:

  • pK1 = 2.19 (α-carboxyl group)
  • pKR = 4.25 (γ-carboxyl group)
  • pK2 = 9.67 (α-amino group)
0 4 8 12 pH 0.0 1.0 2.0 3.0 Equivalents of OH− added pK1 = 2.19 pI = 3.22 pKR = 4.25 pK2 = 9.67 Glu+ Glu2−

The sequential deprotonation pathway from a fully protonated positive state is:

Glu+1 pK1=2.19 Glu0 pKR=4.25 Glu−1 pK2=9.67 Glu−2
  1. At highly acidic pH (pH < 1.0), the molecule exists as Glu+1.
  2. As base is added, the highly acidic α-carboxyl group deprotonates first (pK1 = 2.19), converting the molecule into its zwitterionic form, Glu0.
  3. With further titration, the side chain γ-carboxyl group deprotonates next (pKR = 4.25), converting the neutral zwitterion into a negatively charged anion, Glu−1.

Because the neutral species (Glu0) exists strictly between the first deprotonation (pK1) and the second deprotonation (pKR), the isoelectric point of glutamate is mathematically calculated as the exact average of these two carboxylic pKa values:

pI = (pK1 + pKR) / 2 = (2.19 + 4.25) / 2 = 3.22

Quantitative Titration of a Basic R-Group (e.g., Histidine)

Histidine contains three ionisable groups with the following pKa values:

  • pK1 = 1.82 (α-carboxyl group)
  • pKR = 6.00 (imidazole ring)
  • pK2 = 9.17 (α-amino group)
0 4 8 12 pH 0.0 1.0 2.0 3.0 Equivalents of OH− added pK1 = 1.82 pKR = 6.00 pI = 7.59 pK2 = 9.17 His2+ His

The deprotonation pathway from the fully protonated divalent cation is:

His+2 pK1=1.82 His+1 pKR=6.00 His0 pK2=9.17 His−1
  1. Below pH 1.82, the molecule is a divalent cation, His+2, because both the α-amino and the imidazole ring are fully protonated.
  2. The α-carboxyl deprotonates first (pK1 = 1.82), yielding the monovalent cation His+1.
  3. The imidazole ring deprotonates next (pKR = 6.00), yielding the neutral zwitterion, His0.
  4. Finally, the α-amino group deprotonates at pK2 = 9.17, yielding the anion His−1.

Because the neutral species (His0) exists between the deprotonation of the imidazole ring (pKR) and the α-amino group (pK2), the isoelectric point of histidine is calculated as the average of these two values:

pI = (pKR + pK2) / 2 = (6.00 + 9.17) / 2 = 7.59

Due to its imidazole group having a pKR of 6.00, histidine is the only standard amino acid with a side-chain ionisation constant near physiological pH, enabling it to act as an effective physiological buffer and a key catalytic residue in active sites of enzymes (e.g., the catalytic triad of serine proteases).

6. Peptides, Polypeptides, and Peptide Bond Thermodynamics

Exploring the chemical mechanisms of peptide bond formation, sequence directionality, and quantitative mass calculations in proteins.

Condensation and Dehydration Chemistry

Peptides are polymers of amino acids covalently linked together by peptide bonds (also known as amide bonds). The formation of a peptide bond is a condensation (dehydration) reaction wherein the nucleophilic α-amino group of one amino acid attacks the electrophilic α-carboxyl carbon of an adjacent amino acid, resulting in the elimination of a single water molecule (H2O).

The Condensation Mechanism of Peptide Bond Formation

H3N+–C–C=O R1 H O- + H–N–C–COO- R2 H H H3N+–C–C–N–C–COO- R1 R2 H O H H Peptide Bond + H2O

This reaction is highly endergonic and thermodynamically unfavourable under physiological conditions in water:

ΔG° ≈ +21 kJ/mol

Consequently, to synthesise proteins, living cells must physically couple peptide bond formation directly to the highly exergonic hydrolysis of high-energy nucleoside triphosphates (ATP and GTP) during active ribosomal translation.

Structural Polarisation and Average Residue Weights

Every linear polypeptide has a distinct structural polarity strictly determined by its molecular backbone:

  • The N-terminal (Amino terminus): Positioned at the exact left end of the sequence, featuring a free, unbound α-amino group.
  • The C-terminal (Carboxyl terminus): Positioned at the exact right end of the sequence, featuring a free, unbound α-carboxyl group.

By universal convention, biological peptide sequences are always written, read, and enzymatically synthesised exactly from the N-terminal to the C-terminal.

The average molecular weight of a free standard amino acid monomer is approximately 138 Da (statistically weighted based on typical amino acid abundance in nature). However, when an amino acid is covalently incorporated into a growing polypeptide, a water molecule (Mr = 18) is physically cleaved.

Thus, the average molecular weight of an amino acid residue in a protein is mathematically 110 Da (138 - 18 = 120 Da, which is then analytically adjusted down to 110 Da to accurately account for the significantly higher abundance of smaller amino acids like glycine and alanine in natural structural proteins).

Biophysical Case Studies (Worked Problems)

Problem 1: Fusion Protein Mass Calculation

Scenario: A novel protein X is recombinantly fused to green fluorescent protein (GFP). The gene sequence dictates that Protein X contains precisely 1000 amino acids, and the known molecular mass of free GFP is exactly 27 kDa. What is the total approximate molecular mass of the fused chimeric protein in daltons (Da)?

Solution

Determine the precise molecular mass of the polypeptide chain of Protein X using the established average residue weight:

Mass of Protein X = 1000 residues × 110 Da/residue = 110,000 Da = 110 kDa

Add the molecular mass of the fused GFP domain:

Total Fusion Mass = 110,000 Da + 27,000 Da = 137,000 Da = 137 kDa
Problem 2: Circular Polypeptide Mass Calculation

Scenario: If a free arginine monomer has an exact molecular mass of 174 Da, what is the exact molecular mass (in Daltons) of a circular polymer structurally composed of exactly 38 arginine residues?

Solution

In a standard linear polypeptide containing N residues, there are exactly N-1 peptide bonds, meaning N-1 water molecules are chemically eliminated.
In a circular polypeptide containing N residues, the N-terminal and C-terminal are covalently linked by an additional bridging peptide bond. Therefore, there are exactly N peptide bonds, and exactly N water molecules are eliminated.

Calculate the absolute sum of the free arginine monomers:

Mass of 38 free Arginines = 38 × 174 Da = 6612 Da

Subtract the precise mass of the 38 water molecules eliminated during the covalent circularisation process:

Mass of 38 water molecules = 38 × 18 Da = 684 Da

Molecular Mass of Circular Peptide = 6612 Da - 684 Da = 5928 Da

7. The Stereochemistry of the Peptide Bond

Uncovering the quantum resonance, coplanar rigidity, and cis/trans geometric constraints governing protein backbone folding.

Partial Double-Bond Character and Resonance

In the 1930s and 1940s, Linus Pauling and Robert Corey famously used X-ray crystallography to decisively solve the precise three-dimensional atomic structure of small peptides. They made a fundamental biophysical discovery: the carbon-nitrogen bond linking adjacent amino acid residues is significantly shorter and substantially more rigid than a typical carbon-nitrogen single bond.

Peptide Bond Resonance Stabilization

Structure A O — C – N — H Structure B O- — C = N+ H Resonance Hybrid O(δ-) — C – N(δ+) H

This immense structural rigidity is exclusively due to quantum pi-electron resonance. The lone pair of electrons physically residing on the amide nitrogen is highly delocalised, shifting dynamically into a pi-orbital equally shared with the carbonyl carbon and oxygen. This resonance hybrid physically represents two extreme theoretical structures:

  • Structure A: A single C-N bond paired with a C=O double bond.
  • Structure B: A double C=N bond paired with a C-O single bond.

Consequently, the actual peptide C-N bond permanently possesses approximately 40% double-bond character.

  • A standard single C-N bond is typically measured at 1.49 Å long.
  • A standard double C=N bond is typically measured at 1.27 Å long.
  • The peptide C-N bond is physically measured at exactly 1.33 Å, falling directly and securely between these two theoretical extremes.

The biophysical implications of this profound partial double-bond character are critical:

  1. It completely prevents free geometric rotation around the C-N bond at normal physiological temperatures.
  2. It establishes a permanent, small electric dipole, forcing a partial negative charge (δ-) onto the highly electronegative oxygen atom and a partial positive charge (δ+) onto the nitrogen atom.

The Coplanar Six-Atom Amide Plane

Because the C-N bond physically cannot rotate, the specific group of atoms directly involved in the peptide linkage is sterically constrained to lie entirely within a single, completely rigid, two-dimensional geometric plane, formally termed the amide plane. For any single peptide bond, exactly six atoms mathematically lie in this flat plane:

  1. The α-carbon of the first amino acid (Cα1).
  2. The carbonyl carbon of the first amino acid (C).
  3. The carbonyl oxygen of the first amino acid (O).
  4. The amide nitrogen of the second amino acid (N).
  5. The amide hydrogen of the second amino acid (H).
  6. The α-carbon of the second amino acid (Cα2).

The Six-Atom Coplanar Amide Plane Rigid Amide Plane Cα1 C O N H Cα2

The entire peptide backbone is thus absolutely not a continuously flexible string, but rather a discrete series of entirely rigid, mathematically flat planes uniquely linked by the tetrahedral Cα atoms, which effectively serve as mechanical swivels.

Cis/Trans Isomerism and Steric Clashes

The defined double-bond character of the peptide bond permits two biologically distinct geometric conformations:

  • Trans Configuration: The successive Cα atoms are positioned symmetrically on opposite sides of the rigid peptide bond. The dihedral angle of the peptide bond (ω, omega) is precisely defined as 180°.
  • Cis Configuration: The successive Cα atoms are positioned uniformly on the exact same side of the rigid peptide bond. The dihedral angle ω is precisely defined as .

Geometric Conformational States of the Peptide Bond

Trans Configuration (ω = 180°) Cα1 R1 C O N H Cα2 R2 Cis Configuration (ω = 0°) Cα1 R1 C O N H Cα2 R2

Virtually all structural peptide bonds in correctly folded native proteins occur permanently in the trans configuration. The cis configuration is structurally highly unstable and grossly energetically unfavourable because it geometrically forces the bulky side chains (R groups) of adjacent residues into excessively close physical proximity, generating severe, destabilising steric clashes.

The sole biological exception to this rigid rule is proline. When a peptide bond is covalently formed to a proline residue, the nitrogen atom is mechanically locked within a pyrrolidine ring, which is independently bonded to both the previous carbonyl carbon and its own side chain. Consequently, the steric hindrance is physically comparable in both the cis and trans configurations. While trans logically remains statistically preferred, approximately 10% of proline peptide bonds in native proteins actively adopt the rare cis configuration, which is structurally critical for mechanically forcing tight, hairpin bends and structural loops in complex proteins.

8. Torsion Angles and the Ramachandran Plot

Analyzing the rotational freedom of the polypeptide backbone, steric limitations, and the fundamental map of protein secondary structure.

The Torsion (Dihedral) Angles: Phi (φ) and Psi (ψ)

While the peptide bond C-N is rigid and locked in a plane, the single covalent bonds of the polypeptide backbone on either side of the α-carbon (Cα) are pure single bonds and are free to rotate. This rotational freedom allows the polypeptide chain to fold into various three-dimensional conformations. The conformation of the backbone is completely defined by two torsion (dihedral) angles around each Cα residue:

  • Phi (φ, phi): The angle of rotation around the N-Cα single bond.
  • Psi (ψ, psi): The angle of rotation around the Cα-C single bond.

Polypeptide Backbone Dihedral Angles

C N H φ (Phi) Cα ψ (Psi) C O N

By biochemical convention, both φ and ψ can range from -180° to +180°. When the polypeptide chain is fully extended in a planar conformation, both φ and ψ are geometrically defined as +180° (or -180°). Rotation is defined as positive when looking from the Cα toward the adjacent nitrogen or carbonyl carbon and rotating clockwise.

Biophysical Origins of Steric Hindrance

Although φ and ψ can theoretically adopt any angle between -180° and +180°, the vast majority of these combinations are physically impossible. As the bonds rotate, the non-bonded atoms of the polypeptide backbone (the carbonyl oxygen, amide hydrogen) and the side chains (R groups) collide.

These severe physical collisions occur when the geometric distance between two non-bonded atoms becomes less than the mathematical sum of their van der Waals radii. This intense steric interference severely restricts the actually allowed conformations of the polypeptide backbone to a very small, specific fraction of the total possible φ-ψ map space.

The Architecture of the Ramachandran Plot

In 1963, biophysicist G. N. Ramachandran expertly used strict steric calculations of small peptides to formally map the precisely allowed regions of φ-ψ space. The resulting 2D scatter plot, universally known as the Ramachandran Plot, remains a foundational analytical tool in modern structural biology.

The Standard Ramachandran Plot

β-Sheets (Parallel & Antiparallel) Right-Handed α-Helix Left-Handed α-Helix φ (Phi) Degrees ψ (Psi) Degrees -180 0 +180 +180 0 -180

The plot is systematically divided into distinct topographical regions:

  • Allowed (Shaded/Colored) Regions: Specific combinations of φ and ψ angles that mathematically yield no structural steric clashes. These regions correspond directly to the classic, stable secondary structures found in proteins:
    • Upper Left Quadrant: Contains values corresponding to the flat, extended β-sheets (both antiparallel and parallel conformations) and the structural collagen triple helix.
    • Lower Left Quadrant: Contains values corresponding to the tightly coiled right-handed α-helix.
    • Upper Right Quadrant: A small, isolated allowed island corresponding to the rare left-handed α-helix.
  • Disallowed (White) Regions: Combinations of φ and ψ that are sterically forbidden due to severe atomic collisions.

The Extreme Biophysical Profiles of Glycine and Proline

Two standard amino acids display highly divergent, completely non-standard behaviors on the Ramachandran plot due to their highly unique side-chain physical structures:

1. Glycine (The Hyper-Flexible Conformational Maverick)

Because glycine’s side chain is a single hydrogen atom (-H), it possesses the absolute smallest possible van der Waals volume of any amino acid. As a direct result, glycine experiences minimal to zero steric hindrance during backbone rotation.

On a mapped Ramachandran plot, glycine's allowed regions are exceptionally broad and almost symmetrical across all four quadrants. This high degree of conformational flexibility allows glycine to easily adopt highly unusual dihedral angles, making it a critical, irreplaceable residue in tight structural turns (such as β-turns) and in the uniquely dense structural packaging of the collagen triple helix.

2. Proline (The Conformationally Restricted Ring)

In stark biophysical contrast, proline is the most conformationally restricted amino acid in biology. Because its side-chain carbon aliphatic chain is covalently linked directly back to the active amide nitrogen to form a rigid pyrrolidine ring, the physical rotation specifically around the N-Cα bond is structurally locked.

Reflecting this cyclic geometry, the φ angle of proline is heavily constrained to a highly restricted range of approximately -60° ± 20°. On the Ramachandran plot, proline exhibits an extremely small allowed region. It primarily acts as a severe structural "breaker" of standard α-helices and β-sheets, and is frequently found at the targeted initiation and termination sites of secondary structures, or tightly localized in specialised structural motifs.

In this topic

Scroll to Top