Amino Acids and Peptides
1. Structure, Isomerisation, and Chemical Nature of Amino Acids
The Molecular Architecture of the $\alpha$-Carbon Core
Amino acids serve as the fundamental monomeric building blocks of all proteins. Each $\alpha$-amino acid is organised around a central tetrahedral carbon atom, designated the $\alpha$-carbon ($C_\alpha$). Covalently bound to this asymmetric $C_\alpha$ core are four distinct substituents:
- An acidic carboxyl group ($-\text{COOH}$).
- A basic amino group ($-\text{NH}_2$).
- A unique, variable side chain (the R group) that dictates chemical specificity.
- A solitary hydrogen atom ($-\text{H}$).
α-carboxyl group
COO-
|
H3N+ - Cα - H
| |
α-amino group R
|
Side chain (R group)
The general chemical formula of a standard $\alpha$-amino acid in its non-ionised form is $\text{NH}_2-\text{CHR}-\text{COOH}$. However, because the primary amine and carboxylic acid groups are attached directly to the same $C_\alpha$, these chemical moieties mutually influence one another’s ionisation states, giving rise to unique acid-base dynamics under physiological conditions.
Primary vs. Secondary (Imino) Amino Acids: The Structural Anomaly of Proline
Of the twenty standard, protein-encoded amino acids, nineteen possess a primary amino group ($-\text{NH}_2$) attached to the $C_\alpha$. The sole exception to this structural paradigm is proline.
Proline does not contain a free primary amino group; rather, its side chain is a rigid, five-membered pyrrolidine ring that is covalently bonded to both the $C_\alpha$ and the nitrogen of the amino group. Consequently, proline contains a secondary amino group (an imine, $-\text{NH}-$), making it technically an imino acid.
COO-
/
H2C--Cα--H
/ /
H2C N+H2 <--- Secondary nitrogen locked in ring
\ /
CH2
This cyclic architecture imposes profound steric constraints on the polypeptide backbone, limiting its rotational freedom and significantly influencing the local folding of protein secondary structures.
Non-$\alpha$-Amino Acids: Physiological and Non-Proteinogenic Forms
While proteins are composed exclusively of $\alpha$-amino acids, biologically active, non-proteinogenic amino acids exist in which the amino group is positioned on a carbon atom other than the $C_\alpha$. These are designated as $\beta$-, $\gamma$-, $\delta$-, or $\epsilon$-amino acids, depending on the precise carbon atom hosting the amine moiety relative to the carboxylate terminus.
- $\alpha$-Amino Acid: $\text{H}_3\text{N}^+ – \text{CH}(\text{R}) – \text{COO}^-$
- $\beta$-Amino Acid: $\text{H}_3\text{N}^+ – \text{CH}(\text{R}) – \text{CH}_2 – \text{COO}^-$
- $\gamma$-Amino Acid: $\text{H}_3\text{N}^+ – \text{CH}(\text{R}) – \text{CH}_2 – \text{CH}_2 – \text{COO}^-$
Examples of physiologically critical non-$\alpha$-amino acids include:
- $\beta$-Alanine: A precursor in the biosynthesis of the vitamin pantothenic acid (Vitamin B5) and a constituent of the dipeptide carnosine.
- $\gamma$-Aminobutyric Acid (GABA): A major inhibitory neurotransmitter in the mammalian central nervous system, synthesised via the enzymatic decarboxylation of L-glutamate.
2. Optical Properties and Absolute Configurations
Optical Activity and Polarimetry
All amino acids incorporated into ribosomally synthesised proteins, with the sole exception of glycine, contain at least one asymmetric center—the $C_\alpha$. Glycine is achiral because its R group is a single hydrogen atom, meaning the $C_\alpha$ is bonded to two identical substituents ($-\text{H}$), creating a plane of symmetry.
All other nineteen standard amino acids are chiral molecules; they lack a plane or center of symmetry and are therefore optically active. When a solution of a chiral amino acid is placed in a polarimeter, it rotates the plane of plane-polarised light by a characteristic angle, $\theta$.
[Polarimeter Optical Pathway Schematic]
Light Unpolarised Polariser Plane-polarised Sample Tube Rotated
Source Light Light (Chiral Soln) Light
[O] → -|- [ | ] → | → [ / ] → /
| | /
Optically active molecules are classified by the direction in which they rotate polarised light:
- Dextrorotatory ($d$- or $+$): Rotates the plane of light to the right (clockwise).
- Levorotatory ($l$- or $-$): Rotates the plane of light to the left (counterclockwise).
Biot’s Law and Specific Rotation
The magnitude of optical rotation depends on the concentration of the chiral solute, the temperature, the wavelength of light used, the solvent, and the physical path length of the sample tube. This relationship is quantified by Biot’s Law:
$$ \alpha = [\alpha]^T_\lambda \times C \times l $$
Where:
- $\alpha$ = the observed optical rotation in degrees ($^\circ$).
- $C$ = the concentration of the optically active solute in $\text{g/ml}$.
- $l$ = the physical path length of the sample cell in decimeters ($\text{dm}$, where $1\text{ dm} = 10\text{ cm}$).
- $[\alpha]^T_\lambda$ = the specific rotation of the compound, a fundamental physical constant defined as the observed rotation when plane-polarised light at a specific wavelength $\lambda$ (conventionally the yellow sodium D line, $\lambda = 589\text{ nm}$) passes through a sample cell of $1\text{ dm}$ path length containing a solute concentration of $1\text{ g/ml}$ at temperature $T$ ($^\circ\text{C}$).
The DL Stereochemical System
The absolute configuration of chiral amino acids was historically defined using the DL system, which correlates the spatial arrangement of the substituents around the $C_\alpha$ to the reference three-carbon aldose sugar, glyceraldehyde.
In a standard Fischer projection, the most highly oxidised carbon (the carbonyl/carboxyl group) is oriented at the top of the vertical axis, and the carbon chain extends downwards:
- If the principal functional group (the hydroxyl group in glyceraldehyde, or the amino group in amino acids) is positioned on the left side of the horizontal axis, the molecule is assigned the L-configuration.
- If the functional group is on the right side, the molecule is assigned the D-configuration.
L-Glyceraldehyde D-Glyceraldehyde
CHO CHO
| |
HO - C - H H - C - OH
| |
CH2OH CH2OH
L-Alanine D-Alanine
COO- COO-
| |
H3N+ - C - H H - C - NH3+
| |
CH3 CH3
All chiral amino acids that are ribosomally incorporated into proteins exhibit the L-configuration. D-amino acids are excluded from ribosomal translation, although they occur naturally in non-protein structural biopolymers, such as the peptidoglycan cell walls of Gram-positive and Gram-negative bacteria, and in short peptide antibiotics synthesised via non-ribosomal peptide synthetase (NRPS) pathways.
The Cahn-Ingold-Prelog (RS) Priority Rules
While the DL system is highly convenient for carbohydrates and amino acids, it is an arbitrary comparative system. The modern, systematic framework for defining absolute configuration is the Cahn-Ingold-Prelog (RS) system, which assigns priority to the four substituents bonded to an asymmetric carbon based on their atomic numbers.
The basic rules for priority determination are:
- Atoms directly bonded to the chiral center are ranked by decreasing atomic number ($-\text{S} > -\text{O} > -\text{N} > -\text{C} > -\text{H}$).
- If two or more attached atoms are identical, priority is determined by comparing the atomic numbers of the next atoms in their respective chains.
- Double and triple bonds are treated as if they were bonded to two or three of those atoms via single bonds (e.g., a carboxyl group $-\text{COO}^-$ is treated as a carbon bonded to three oxygen atoms).
The priority sequence for standard amino acid substituents is typically:
$$ -\text{SH} > -\text{OH} > -\text{NH}_2 > -\text{COOH} > -\text{CH}_2\text{OH} > -\text{CH}_3 > -\text{H} $$
[RS System Three-Dimensional Projection]
(1) [NH3+]
|
Cα
/ \ \
(3) [CH3] / \ \ (4) [H] (points away)
/ \
(2) [COO-]
Direction (1 → 2 → 3): Counterclockwise (S)
To assign the configuration:
- The molecule is viewed with the lowest priority group (typically hydrogen, priority 4) pointing directly away from the observer.
- The remaining three substituents (priorities 1, 2, and 3) are traced in descending order:
- If the path from 1 $\rightarrow$ 2 $\rightarrow$ 3 runs clockwise, the configuration is R (rectus, right).
- If the path runs counterclockwise, the configuration is S (sinister, left).
Under this system, almost all L-amino acids found in proteins exhibit the S configuration at the $C_\alpha$ (for example, L-alanine is (S)-alanine). The single, critical exception is L-cysteine. Because the sulfur atom in the thiol side chain ($-\text{CH}_2\text{SH}$) has a higher atomic number than the oxygen atoms of the carboxyl group ($-\text{COO}^-$), the priority rankings shift:
- $-\text{NH}_3^+$ (Priority 1)
- $-\text{CH}_2\text{SH}$ (Priority 2)
- $-\text{COO}^-$ (Priority 3)
- $-\text{H}$ (Priority 4)
Traced from 1 $\rightarrow$ 2 $\rightarrow$ 3, the path runs clockwise, designating L-cysteine as (R)-cysteine.
3. Taxonomy and Structural Classification of Amino Acids
Proteinogenic (Standard) vs. Non-Standard Amino Acids
Of the hundreds of amino acids existing in nature, only 22 are classified as standard (proteinogenic) amino acids, meaning they are specified by genetic codons and incorporated into nascent polypeptide chains during translation.
Non-standard amino acids are those that are not directly incorporated via standard ribosomal translation. Instead, they are generated through highly regulated post-translational modifications (PTMs) of standard residues already embedded within a polypeptide chain, or they function as key metabolic intermediates.
- 4-Hydroxyproline: Generated by the post-translation modification of proline residues by prolyl 4-hydroxylase; essential for the thermal stability of the collagen triple helix.
- 5-Hydroxylysine: Synthesised by lysyl hydroxylase; critical for the covalent cross-linking of collagen fibers.
- Desmosine: A complex, tetra-functional amino acid formed by the condensation of four lysine side chains, providing the elastic, rubber-like properties of elastin.
- $\gamma$-Carboxyglutamate: Synthesised via a Vitamin K-dependent carboxylation of glutamate; contains two adjacent carboxyl groups in its side chain, enabling high-affinity binding of calcium ions ($\text{Ca}^{2+}$) in clotting factors like prothrombin.
- N-Formylmethionine (fMet): The initiator amino acid in prokaryotic translation and eukaryotic mitochondrial/chloroplast organellar protein synthesis.
Non-Standard Amino Acid Structures:
4-Hydroxyproline: γ-Carboxyglutamate:
HO-CH---CH2 COO- COO-
| | \ /
H2C CH-COO- CH
\ / |
N+H2 CH2
|
H3N+ - C - H
|
COO-
The 21st and 22nd Standard Amino Acids
The structural taxonomy of proteinogenic amino acids is extended by two rare, specialised residues that are co-translationally incorporated in response to specific stop codons through unique transfer RNA (tRNA) reprogramming:
- Selenocysteine (Sec, U): Known as the “21st amino acid,” its structure is a homologue of cysteine in which the sulfur atom is replaced by a selenium atom ($-\text{CH}_2\text{SeH}$). It is incorporated during translation in response to the stop codon UGA. This requires a specific downstream mRNA stem-loop structure called the SECIS (Selenocysteine Insertion Sequence) element, which recruits a dedicated translation elongation factor (SelB) and a unique tRNA ($\text{tRNA}^{\text{Sec}}$). Selenocysteine is a catalytic residue in vital redox enzymes, including glutathione peroxidase and formate dehydrogenase.
- Pyrrollysine (Pyl, O): The “22nd amino acid,” found in methanogenic archaea and certain anaerobic bacteria. It consists of a lysine residue covalently linked to a pyrroline ring. It is incorporated at the stop codon UAG via a specialised pyrrollysyl-tRNA synthetase and a dedicated $\text{tRNA}^{\text{Pyl}}$, and is critical for the catalytic function of methyltransferases in methanogenesis.
Selenocysteine (Sec) Pyrrollysine (Pyl)
COO- COO-
| |
H3N+ - C - H H3N+ - C - H
| |
CH2 (CH2)4
| |
SeH NH
|
C = O
/
HC == N
/ \
H3C-CH--CH2
Structural Categorisation by R-Group Chemistry (at Physiological pH 7.4)
The physical, chemical, and structural roles of the 20 standard amino acids are determined entirely by the properties of their side chains (R groups). At physiological pH (~7.4), they are classified into four primary categories:
1. Non-Polar, Hydrophobic Side Chains (9 Amino Acids)
These side chains lack polar functional groups, possess low dielectric constants, and tend to cluster within the interior of folded proteins to avoid contact with the aqueous solvent, driving the hydrophobic collapse of proteins.
- Glycine (Gly, G): R = $-\text{H}$. The smallest, most flexible amino acid; lacks chirality.
- Alanine (Ala, A): R = $-\text{CH}_3$. A simple methyl group.
- Valine (Val, V): R = $-\text{CH}(\text{CH}_3)_2$. A branched-chain aliphatic hydrocarbon.
- Leucine (Leu, L): R = $-\text{CH}_2-\text{CH}(\text{CH}_3)_2$. Isomeric with isoleucine.
- Isoleucine (Ile, I): R = $-\text{CH}(\text{CH}_3)-\text{CH}_2\text{CH}_3$. Contains a second chiral center at the $\beta$-carbon.
- Proline (Pro, P): R = $-\text{CH}_2-\text{CH}_2-\text{CH}_2-$, forming a secondary imine pyrrolidine ring.
- Methionine (Met, M): R = $-\text{CH}_2-\text{CH}_2-\text{S}-\text{CH}_3$. An ether-linked thioether.
- Phenylalanine (Phe, F): R = $-\text{CH}_2-\text{C}_6\text{H}_5$. A highly hydrophobic phenyl ring.
- Tryptophan (Trp, W): R = $-\text{CH}_2-\text{indole ring}$. A bulky, bicyclic aromatic system.
2. Polar, Uncharged (Hydrophilic) Side Chains (6 Amino Acids)
These side chains possess electronegative heteroatoms (oxygen, nitrogen, or sulfur) that can engage in hydrogen bonding with water or other polar residues, but they do not carry a net charge at neutral pH.
- Serine (Ser, S): R = $-\text{CH}_2\text{OH}$. Contains a primary hydroxyl group.
- Threonine (Thr, T): R = $-\text{CH}(\text{OH})-\text{CH}_3$. Contains a secondary hydroxyl and a second chiral center.
- Cysteine (Cys, C): R = $-\text{CH}_2\text{SH}$. Contains a highly reactive thiol (sulfhydryl) group, capable of forming covalent disulfide bonds (bridges) under oxidising conditions.
- Asparagine (Asn, N): R = $-\text{CH}_2-\text{CONH}_2$. The amide derivative of aspartate.
- Glutamine (Gln, Q): R = $-\text{CH}_2-\text{CH}_2-\text{CONH}_2$. The amide derivative of glutamate.
- Tyrosine (Tyr, Y): R = $-\text{CH}_2-\text{phenol ring}$. An amphipathic aromatic residue with a weakly acidic hydroxyl group.
3. Positively Charged (Basic) Polar Side Chains (3 Amino Acids)
These side chains contain nitrogenous bases that are protonated and carry a net positive charge at physiological pH.
- Lysine (Lys, K): R = $-\text{CH}_2-\text{CH}_2-\text{CH}_2-\text{CH}_2-\text{NH}_3^+$. Terminated by a flexible primary $\epsilon$-amino group.
- Arginine (Arg, R): R = $-\text{CH}_2-\text{CH}_2-\text{CH}_2-\text{NH}-\text{C}(\text{NH}_2)=\text{NH}_2^+$. Terminated by a highly basic guanidinium group, which remains protonated across almost the entire physiological pH spectrum.
- Histidine (His, H): R = $-\text{CH}_2-\text{imidazole ring}$. Contains an weakly basic imidazole group. With a $\text{p}K_a$ close to neutral pH, it can readily transition between protonated (positive) and deprotonated (neutral) states, serving as an essential proton donor/acceptor in enzyme-catalysed reactions.
4. Negatively Charged (Acidic) Polar Side Chains (2 Amino Acids)
These side chains contain carboxyl groups that are fully deprotonated and carry a net negative charge at physiological pH.
- Aspartate (Asp, D): R = $-\text{CH}_2-\text{COO}^-$. A $\beta$-carboxyl group.
- Glutamate (Glu, E): R = $-\text{CH}_2-\text{CH}_2-\text{COO}^-$. A $\gamma$-carboxyl group.
4. Ultraviolet Spectroscopy of Aromatic Amino Acids
The Physics of UV Absorption in Proteins
The three aromatic amino acids—phenylalanine, tyrosine, and tryptophan—contain conjugated $\pi$-electron systems that can undergo electronic transitions upon absorbing ultraviolet (UV) radiation. This property is exploited for the non-destructive detection, characterisation, and quantitative measurement of proteins in biochemical solutions.
Absorbance
^
| _ (Tryptophan, Trp, λmax = 279.8 nm)
| / \
| / \ _ (Tyrosine, Tyr, λmax = 274.6 nm)
| / \ / \
| / \ / \ .. (Phenylalanine, Phe, λmax = 257.4 nm)
| / \/ \ :
+-----------------------------------> Wavelength (nm)
240 250 260 270 280
The specific spectral properties of the three aromatic residues are:
- Phenylalanine: Exhibits a weak absorption maximum at $257.4\text{ nm}$ due to its single phenyl ring. Its molar extinction coefficient ($\epsilon$) is exceptionally low, making it a negligible contributor to the overall UV absorbance of typical proteins.
- Tyrosine: Features a strong absorption peak at $274.6\text{ nm}$, originating from its phenolic ring. At high pH (>10), deprotonation of the phenolic hydroxyl group shifts the absorption maximum bathochromically (red-shift) to longer wavelengths and increases the extinction coefficient.
- Tryptophan: Displays the most intense absorption profile, with a major peak at $279.8\text{ nm}$. This dominant absorption is due to its highly conjugated, bicyclic heterocyclic indole ring.
The 280 nm Spectroscopic Assay
In laboratory biochemistry, protein concentration is routinely measured by monitoring absorbance at $280\text{ nm}$. This wavelength represents a sweet spot where the absorption of tryptophan and tyrosine dominates the spectrum:
- Phenylalanine does not absorb significantly at $280\text{ nm}$.
- Tryptophan’s absorbance at $280\text{ nm}$ is approximately four times greater than that of tyrosine on a molar basis, meaning the total $A_{280}$ of a protein is heavily determined by its relative tryptophan content.
- The protein concentration is calculated from the $A_{280}$ value using the Beer-Lambert Law:
$$ A = \epsilon \times c \times b $$
Where:
- $A$ = measured absorbance (dimensionless).
- $\epsilon$ = molar extinction coefficient ($\text{M}^{-1}\text{ cm}^{-1}$), which can be calculated theoretically based on the number of tryptophan and tyrosine residues: $\epsilon_{280} \approx (n_{\text{Trp}} \times 5500) + (n_{\text{Tyr}} \times 1490)$.
- $c$ = molar concentration of the protein ($\text{M}$).
- $b$ = optical path length of the cuvette (typically $1\text{ cm}$).
5. Acid-Base Chemistry and Titration Dynamics
Amphiprotic (Amphoteric) Nature and Zwitterions
Amino acids are classic examples of amphiprotic (or amphoteric) substances; they contain both proton-donating acidic groups ($-\text{COOH}$) and proton-accepting basic groups ($-\text{NH}_2$), allowing them to act as either acids or bases depending on the pH of the surrounding solution.
At physiological pH (~7.4), the carboxylic acid group ($-\text{COOH}$) is fully deprotonated to its conjugate base form ($-\text{COO}^-$), while the primary amino group ($-\text{NH}_2$) is protonated to its conjugate acid form ($-\text{NH}_3^+$). This dipolar ion species, which carries equal and opposite charges on different atoms and has a net charge of zero, is called a zwitterion (from the German Zwitter, meaning “hybrid”).
Low pH (highly acidic) Physiological pH (~7.4) High pH (highly basic)
COOH COO- COO-
| | |
H3N+ - C - H H3N+ - C - H H2N - C - H
| | |
R R R
Net Charge: +1 Net Charge: 0 Net Charge: -1
Quantitative Titration of a Non-Ionisable R-Group (e.g., Alanine)
For amino acids with simple aliphatic or uncharged side chains, there are only two ionisable groups: the $\alpha$-carboxyl group and the $\alpha$-amino group. If we titrate the fully protonated cationic form of alanine (at pH < 1.0) with a strong base (such as $\text{NaOH}$), the proton-release occurs in a stepwise, two-stage process:
pH
12 + __ (Ala-)
| _--
10 + _-- [pK2 = 9.69] <-- Amino Buffer region
| _--
8 + /-
| /
6 +-------------------------------* [pI = 6.01] <-- Isoelectric Point (Zwitterion)
| /
4 + /-
| _--
2 + _-- [pK1 = 2.34] <-- Carboxyl Buffer region
| _-
0 +___________________/ (Ala+)
0.0 1.0 2.0
Equivalents of OH- added
- The First Dissociation Stage (Deprotonation of the $\alpha$-Carboxyl Group): At highly acidic pH, the dominant species is the fully protonated cation, $\text{Ala}^+$ (net charge $+1$). As hydroxide ions ($\text{OH}^-$) are added, the highly acidic proton of the carboxyl group is released first:
$$ \text{Ala}^+ + \text{OH}^- \rightleftharpoons \text{Ala}^0 + \text{H}_2\text{O} $$
The midpoint of this first titration stage is reached when exactly $0.5$ equivalents of base have been added. At this point, the concentrations of the cation ($\text{Ala}^+$) and the zwitterion ($\text{Ala}^0$) are equal ($[\text{Ala}^+] = [\text{Ala}^0]$). The pH at this midpoint is defined as the $\text{p}K_1$ (or $\text{p}K_a$ of the carboxyl group), which for alanine is $2.34$. This region exhibits strong buffering capacity, as described by the Henderson-Hasselbalch equation:
$$ \text{pH} = \text{p}K_1 + \log \frac{[\text{Ala}^0]}{[\text{Ala}^+]} $$
- The Second Dissociation Stage (Deprotonation of the $\alpha$-Amino Group): As the titration continues past $\text{pH} = 6.0$, the zwitterion ($\text{Ala}^0$) is the dominant species. With further addition of base, the weaker acid (the protonated primary amino group, $-\text{NH}_3^+$) begins to lose its proton:
$$ \text{Ala}^0 + \text{OH}^- \rightleftharpoons \text{Ala}^- + \text{H}_2\text{O} $$
The midpoint of this second stage occurs when $1.5$ equivalents of base have been added. Here, the concentration of the zwitterion ($\text{Ala}^0$) equals that of the anionic species ($\text{Ala}^-$). The pH at this midpoint is defined as $\text{p}K_2$ (or $\text{p}K_a$ of the amino group), which for alanine is $9.69$.
Calculation of the Isoelectric Point ($\text{pI}$)
The isoelectric point ($\text{pI}$) is defined as the precise pH at which the net electric charge of the amino acid population is exactly zero. At this pH, the molecule is electrophoretically immobile (it will not migrate in an electric field) and displays its minimum solubility in water, as it lacks a net repulsive charge.
For amino acids with non-ionisable side chains, the zwitterion is the intermediate species flanking the protonation states of the carboxyl and amino groups. Therefore, the $\text{pI}$ is mathematically derived as the simple arithmetic mean of the two $\text{p}K_a$ values:
$$ \text{pI} = \frac{\text{p}K_1 + \text{p}K_2}{2} $$
For alanine:
$$ \text{pI} = \frac{2.34 + 9.69}{2} = 6.015 \approx 6.02 $$
Quantitative Titration of an Acidic R-Group (e.g., Glutamate)
Amino acids with an acidic side chain contain three ionisable groups: the $\alpha$-carboxyl group, the $\gamma$-carboxyl group of the side chain, and the $\alpha$-amino group. The respective $\text{p}K_a$ values for glutamic acid are:
- $\text{p}K_1 = 2.19$ ($\alpha$-carboxyl group)
- $\text{p}K_{\text{R}} = 4.25$ ($\gamma$-carboxyl group)
- $\text{p}K_2 = 9.67$ ($\alpha$-amino group)
pH
12 + __ (Glu2-)
| _--
10 + _-- [pK2 = 9.67]
| _--
8 + /-
| /
6 + /
| /
4 + _-------* [pKR = 4.25]
| _-
3 + -* [pI = 3.22]
| _-
2 + _-* [pK1 = 2.19]
| _-
0 +_____________/ (Glu+)
0.0 1.0 2.0 3.0
Equivalents of OH- added
The sequential deprotonation pathway from a fully protonated positive state is:
$$ \text{Glu}^{+1} \xrightarrow[\text{p}K_1 = 2.19]{-\text{H}^+} \text{Glu}^0 \xrightarrow[\text{p}K_{\text{R}} = 4.25]{-\text{H}^+} \text{Glu}^{-1} \xrightarrow[\text{p}K_2 = 9.67]{-\text{H}^+} \text{Glu}^{-2} $$
- At highly acidic pH (pH < 1.0), the molecule exists as $\text{Glu}^{+1}$.
- As base is added, the highly acidic $\alpha$-carboxyl group deprotonates first ($\text{p}K_1 = 2.19$), converting the molecule into its zwitterionic form, $\text{Glu}^0$.
- With further titration, the side chain $\gamma$-carboxyl group deprotonates next ($\text{p}K_{\text{R}} = 4.25$), converting the neutral zwitterion into a negatively charged anion, $\text{Glu}^{-1}$.
- Because the neutral species ($\text{Glu}^0$) exists between the first deprotonation ($\text{p}K_1$) and the second deprotonation ($\text{p}K_{\text{R}}$), the isoelectric point of glutamate is calculated as the average of these two carboxylic $\text{p}K_a$ values:
$$ \text{pI} = \frac{\text{p}K_1 + \text{p}K_{\text{R}}}{2} = \frac{2.19 + 4.25}{2} = 3.22 $$
Quantitative Titration of a Basic R-Group (e.g., Histidine)
Histidine contains three ionisable groups with the following $\text{p}K_a$ values:
- $\text{p}K_1 = 1.82$ ($\alpha$-carboxyl group)
- $\text{p}K_{\text{R}} = 6.00$ (imidazole ring)
- $\text{p}K_2 = 9.17$ ($\alpha$-amino group)
pH
12 + __ (His-)
| _--
10 + _-- [pK2 = 9.17]
| _--
8 + /-
| /
7.6 +---------------------------------* [pI = 7.59]
| /
6 + _-------* [pKR = 6.00]
| _-
4 + /-
| /
2 + _-* [pK1 = 1.82]
| _-
0 +______________/ (His2+)
0.0 1.0 2.0 3.0
Equivalents of OH- added
The deprotonation pathway from the fully protonated divalent cation is:
$$ \text{His}^{+2} \xrightarrow[\text{p}K_1 = 1.82]{-\text{H}^+} \text{His}^{+1} \xrightarrow[\text{p}K_{\text{R}} = 6.00]{-\text{H}^+} \text{His}^0 \xrightarrow[\text{p}K_2 = 9.17]{-\text{H}^+} \text{His}^{-1} $$
- Below pH 1.82, the molecule is a divalent cation, $\text{His}^{+2}$, because both the $\alpha$-amino and the imidazole ring are fully protonated.
- The $\alpha$-carboxyl deprotonates first ($\text{p}K_1 = 1.82$), yielding the monovalent cation $\text{His}^{+1}$.
- The imidazole ring deprotonates next ($\text{p}K_{\text{R}} = 6.00$), yielding the neutral zwitterion, $\text{His}^0$.
- Finally, the $\alpha$-amino group deprotonates at $\text{p}K_2 = 9.17$, yielding the anion $\text{His}^{-1}$.
- Because the neutral species ($\text{His}^0$) exists between the deprotonation of the imidazole ring ($\text{p}K_{\text{R}}$) and the $\alpha$-amino group ($\text{p}K_2$), the isoelectric point of histidine is calculated as the average of these two values:
$$ \text{pI} = \frac{\text{p}K_{\text{R}} + \text{p}K_2}{2} = \frac{6.00 + 9.17}{2} = 7.59 $$
Due to its imidazole group having a $\text{p}K_{\text{R}}$ of $6.00$, histidine is the only standard amino acid with a side-chain ionisation constant near physiological pH, enabling it to act as an effective physiological buffer and a key catalytic residue in active sites of enzymes (e.g., the catalytic triad of serine proteases).
6. Peptides, Polypeptides, and Peptide Bond Thermodynamics
Condensation and Dehydration Chemistry
Peptides are polymers of amino acids covalently linked together by peptide bonds (also known as amide bonds). The formation of a peptide bond is a condensation (dehydration) reaction wherein the nucleophilic $\alpha$-amino group of one amino acid attacks the electrophilic $\alpha$-carboxyl carbon of an adjacent amino acid, resulting in the elimination of a single water molecule ($\text{H}_2\text{O}$).
R1 R2
| |
H3N+-C-C=O + H-N-C-COO- → H3N+-C-C-N-C-COO- + H2O
| | | | | | | |
H O- H H H O H H
\_______/
Peptide Bond
This reaction is highly endergonic and thermodynamically unfavourable under physiological conditions in water:
$$ \Delta G^\circ \approx +21\text{ kJ/mol} $$
Consequently, to synthesise proteins, living cells must couple peptide bond formation to the hydrolysis of high-energy nucleoside triphosphates (ATP and GTP) during translation.
Structural Polarisation and Average Residue Weights
Every linear polypeptide has a distinct structural polarity determined by its molecular backbone:
- The N-terminal (Amino terminus): Positioned at the left end of the sequence, featuring a free $\alpha$-amino group.
- The C-terminal (Carboxyl terminus): Positioned at the right end of the sequence, featuring a free $\alpha$-carboxyl group.
By convention, peptide sequences are always written, read, and synthesised from the N-terminal to the C-terminal.
The average molecular weight of a free standard amino acid is approximately $138\text{ Da}$ (weighted based on typical amino acid abundance in nature). However, when an amino acid is incorporated into a polypeptide, a water molecule ($M_r = 18$) is cleaved. Thus, the average molecular weight of an amino acid residue in a protein is $110\text{ Da}$ ($138 – 18 = 120\text{ Da}$, adjusted down to $110\text{ Da}$ to account for the higher abundance of smaller amino acids like glycine and alanine in structural proteins).
Biophysical Case Studies (Worked Problems)
Problem 1: Fusion Protein Mass Calculation
- Scenario: A novel protein X is recombinantly fused to green fluorescent protein (GFP). The gene sequence dictates that Protein X contains $1000$ amino acids, and the molecular mass of free GFP is $27\text{ kDa}$. What is the total approximate molecular mass of the fused chimeric protein in daltons (Da)?
- Solution:
- Determine the molecular mass of the polypeptide chain of Protein X using the average residue weight:
$$ \text{Mass of Protein X} = 1000 \text{ residues} \times 110 \text{ Da/residue} = 110,000 \text{ Da} = 110 \text{ kDa} $$ - Add the molecular mass of the fused GFP:
$$ \text{Total Fusion Mass} = 110,000 \text{ Da} + 27,000 \text{ Da} = 137,000 \text{ Da} = 137 \text{ kDa} $$
- Determine the molecular mass of the polypeptide chain of Protein X using the average residue weight:
Problem 2: Circular Polypeptide Mass Calculation
- Scenario: If a free arginine monomer has a molecular mass of $174\text{ Da}$, what is the exact molecular mass (in Daltons) of a circular polymer composed of $38$ arginine residues?
- Solution:
- In a linear polypeptide containing $N$ residues, there are $N-1$ peptide bonds, meaning $N-1$ water molecules are eliminated.
- In a circular polypeptide containing $N$ residues, the N-terminal and C-terminal are covalently linked by an additional peptide bond. Therefore, there are exactly $N$ peptide bonds, and exactly $N$ water molecules are eliminated.
- Calculate the sum of the free arginine monomers:
$$ \text{Mass of 38 free Arginines} = 38 \times 174 \text{ Da} = 6612 \text{ Da} $$ - Subtract the mass of $38$ water molecules eliminated during the circularisation process:
$$ \text{Mass of 38 water molecules} = 38 \times 18 \text{ Da} = 684 \text{ Da} $$
$$ \text{Molecular Mass of Circular Peptide} = 6612 \text{ Da} – 684 \text{ Da} = 5928 \text{ Da} $$
7. The Stereochemistry of the Peptide Bond
Partial Double-Bond Character and Resonance
In the 1930s and 1940s, Linus Pauling and Robert Corey used X-ray crystallography to solve the precise three-dimensional structure of small peptides. They made a fundamental discovery: the carbon-nitrogen bond linking adjacent amino acid residues is significantly shorter and more rigid than a typical carbon-nitrogen single bond.
Structure A Structure B
O O-
|| |
-- C - N -- ⇔ -- C = N+ --
| |
H H
Resonance Hybrid
O(δ-)
|
-- C --- N(δ+) --
.
|
H
This structural rigidity is due to pi-electron resonance. The lone pair of electrons on the amide nitrogen is delocalised, shifting into a pi-orbital shared with the carbonyl carbon and oxygen. This resonance hybrid has two extreme structures:
- Structure A: A single C-N bond with a C=O double bond.
- Structure B: A double C=N bond with a C-O single bond.
Consequently, the peptide C-N bond has approximately 40% double-bond character.
- A standard single C-N bond is typically $1.49\text{ \AA}$ long.
- A standard double C=N bond is typically $1.27\text{ \AA}$ long.
- The peptide C-N bond is measured at $1.33\text{ \AA}$, falling directly between these values.
The implications of this partial double-bond character are profound:
- It prevents free rotation around the C-N bond at physiological temperatures.
- It sets up a small electric dipole, with a partial negative charge ($\delta-$) on the highly electronegative oxygen atom and a partial positive charge ($\delta+$) on the nitrogen atom.
The Coplanar Six-Atom Amide Plane
Because the C-N bond cannot rotate, the group of atoms directly involved in the peptide linkage is constrained to lie in a single, rigid, two-dimensional geometric plane, termed the amide plane. For any peptide bond, exactly six atoms lie in this plane:
- The $\alpha$-carbon of the first amino acid ($C_{\alpha1}$).
- The carbonyl carbon of the first amino acid ($C$).
- The carbonyl oxygen of the first amino acid ($O$).
- The amide nitrogen of the second amino acid ($N$).
- The amide hydrogen of the second amino acid ($H$).
- The $\alpha$-carbon of the second amino acid ($C_{\alpha2}$).
[The Six-Atom Coplanar Amide Plane]
O
||
- Cα1 - C --------------- N - Cα2 -
|
H
\____________________________________________/
Rigid Amide Plane
The peptide backbone is thus not a continuously flexible string, but rather a series of rigid, flat planes linked by the tetrahedral $C_\alpha$ atoms, which serve as swivels.
Cis/Trans Isomerism and Steric Clashes
The double-bond character of the peptide bond permits two distinct geometric conformations:
- Trans Configuration: The successive $C_\alpha$ atoms are positioned on opposite sides of the peptide bond. The dihedral angle of the peptide bond ($\omega$, omega) is defined as $180^\circ$.
- Cis Configuration: The successive $C_\alpha$ atoms are positioned on the same side of the peptide bond. The dihedral angle $\omega$ is defined as $0^\circ$.
Trans Configuration (ω = 180°)
R1 H
\ /
Cα1 - C == N - Cα2
/ \
O R2
Cis Configuration (ω = 0°)
R1 R2
\ /
Cα1 - C == N - Cα2
/ \
O H
Virtually all peptide bonds in native proteins occur in the trans configuration. The cis configuration is highly unstable and energetically unfavourable because it forces the bulky side chains (R groups) of adjacent residues into close physical proximity, generating severe steric clashes.
The sole exception to this rule is proline. When a peptide bond is formed to a proline residue, the nitrogen atom is locked in a pyrrolidine ring, which is bonded to both the previous carbonyl carbon and its own side chain. Consequently, the steric hindrance is comparable in both the cis and trans configurations. While trans remains preferred, approximately 10% of proline peptide bonds in native proteins adopt the cis configuration, which is critical for making tight bends and loops in structural proteins.
8. Torsion Angles and the Ramachandran Plot
The Torsion (Dihedral) Angles: Phi ($\phi$) and Psi ($\psi$)
While the peptide bond C-N is rigid and locked in a plane, the single covalent bonds of the polypeptide backbone on either side of the $C_\alpha$ are pure single bonds and are free to rotate. This rotational freedom allows the polypeptide chain to fold into various three-dimensional conformations. The conformation of the backbone is completely defined by two torsion (dihedral) angles around each $C_\alpha$ residue:
- C - [N - Cα] - [Cα - C] - N -
| |
Phi (φ) Psi (ψ)
- Phi ($\phi$, phi): The angle of rotation around the $\text{N}-\text{C}_\alpha$ single bond.
- Psi ($\psi$, psi): The angle of rotation around the $\text{C}_\alpha-\text{C}$ single bond.
By biochemical convention, both $\phi$ and $\psi$ can range from $-180^\circ$ to $+180^\circ$. When the polypeptide chain is fully extended in a planar conformation, both $\phi$ and $\psi$ are defined as $+180^\circ$ (or $-180^\circ$). Rotation is defined as positive when looking from the $C_\alpha$ toward the adjacent nitrogen or carbonyl carbon and rotating clockwise.
Biophysical Origins of Steric Hindrance
Although $\phi$ and $\psi$ can theoretically adopt any angle between $-180^\circ$ and $+180^\circ$, the vast majority of these combinations are physically impossible. As the bonds rotate, the non-bonded atoms of the polypeptide backbone (the carbonyl oxygen, amide hydrogen) and the side chains (R groups) collide.
These collisions occur when the distance between two non-bonded atoms is less than the sum of their van der Waals radii. This steric interference restricts the allowed conformations of the polypeptide backbone to a small fraction of the total possible $\phi-\psi$ space.
The Architecture of the Ramachandran Plot
In 1963, G. N. Ramachandran used steric calculations of small peptides to map the allowed regions of $\phi-\psi$ space. The resulting 2D scatter plot, known as the Ramachandran Plot, is a foundational tool in structural biology.
ψ (degrees)
+180 +-----------------------------------------+
| [β-Sheets] |
| (Antiparallel |
| & Parallel) |
| :--- |
0 | [Left-Handed] |
| α-Helix |
| :--- |
| [Right-Handed] |
| α-Helix |
-180 +-----------------------------------------+
-180 0 +180
φ (degrees)
The plot is divided into distinct topographical regions:
- Allowed (Shaded/Black) Regions: Combinations of $\phi$ and $\psi$ angles that yield no steric clashes. These regions correspond directly to the classic, stable secondary structures found in proteins:
- Upper Left Quadrant: Contains values corresponding to the flat, extended $\beta$-sheets (both antiparallel and parallel conformations) and the structural collagen triple helix.
- Lower Left Quadrant: Contains values corresponding to the tightly coiled right-handed $\alpha$-helix.
- Upper Right Quadrant: A small, isolated allowed island corresponding to the rare left-handed $\alpha$-helix.
- Disallowed (White) Regions: Combinations of $\phi$ and $\psi$ that are sterically forbidden due to atomic collisions.
The Extreme Biophysical Profiles of Glycine and Proline
Two standard amino acids display highly divergent, non-standard behaviors on the Ramachandran plot due to their unique side-chain structures:
1. Glycine (The Hyper-Flexible Conformational Maverick)
Because glycine’s side chain is a single hydrogen atom ($-\text{H}$), it has the smallest possible van der Waals volume of any amino acid. As a result, glycine experiences minimal steric hindrance during backbone rotation.
On a Ramachandran plot, glycine’s allowed regions are exceptionally broad and symmetrical across all four quadrants. This high degree of conformational flexibility allows glycine to adopt unusual dihedral angles, making it a critical residue in tight structural turns (such as $\beta$-turns) and in the dense structural packaging of the collagen triple helix.
2. Proline (The Conformationally Restricted Ring)
In contrast, proline is the most conformationally restricted amino acid. Because its side-chain carbon chain is covalently linked back to the amide nitrogen to form a pyrrolidine ring, the rotation around the $\text{N}-\text{C}_\alpha$ bond is physically locked.
Reflecting this cyclic geometry, the $\phi$ angle of proline is constrained to a highly restricted range of approximately $-60^\circ \pm 20^\circ$. On the Ramachandran plot, proline exhibits an extremely small allowed region. It acts as a structural “breaker” of standard $\alpha$-helices and $\beta$-sheets, and is frequently found at the initiation and termination sites of secondary structures, or in specialised structural motifs.
