Protein Folding

The Physics and Thermodynamics of Protein Folding

The Physics and Thermodynamics of Protein Folding

An exploration of the thermodynamic drivers, free energy states, the hydrophobic effect, and enthalpic contributions stabilizing the native conformation.

Protein folding is the highly coordinated physical process by which a linear, unstructured polypeptide chain folds into its unique, biologically active, and stable three-dimensional conformation, termed the native state.

Protein Folding Transition Unstructured Polypeptide High conformational entropy High free energy (G) Thermodynamically unstable Spontaneous ΔG < 0 Native Conformation Low conformational entropy Minimum free energy (G) Thermodynamically stable

The native conformation is dictated entirely by the primary amino acid sequence. Failure of a polypeptide to adopt its native state results in non-functional, inactive, or toxic misfolded species. Although protein folding occurs spontaneously in vivo and in vitro for many proteins, the intracellular environment presents extreme macromolecular crowding, requiring molecular chaperones to prevent off-pathway aggregation.

1.1 Thermodynamic Drivers of Protein Folding

The transition of an unfolded polypeptide to a folded macromolecule is thermodynamically spontaneous, meaning the change in Gibbs free energy (ΔGfolding) is negative:

ΔGfolding = ΔHfolding - TΔSfolding < 0

At first glance, the folding process appears to violate the Second Law of Thermodynamics, as organizing a highly flexible, random-coil polymer into a single, well-defined conformation represents a massive decrease in the conformational entropy of the polypeptide chain (ΔSpolypeptide ≪ 0). This unfavourable entropic penalty is overcome by two major thermodynamic factors:

The Hydrophobic Effect (Entropic Driver)

This is the predominant thermodynamic force driving protein folding. In an unfolded state, hydrophobic side chains (such as Leu, Ile, Val, Phe, Trp) are exposed to the bulk aqueous solvent. This exposure forces surrounding water molecules to organize into highly structured, rigid, cage-like structures (clathrates) to maximize hydrogen bonding. This represents a significant local decrease in solvent entropy. Upon spontaneous collapse of the hydrophobic residues into the interior core of the protein, these ordered water molecules are released back into the bulk solvent, causing a massive increase in solvent entropy (ΔSsolvent ≫ 0). The net entropy change of the system (polypeptide + solvent) is highly positive:

ΔStotal = ΔSpolypeptide + ΔSsolvent > 0

Enthalpic Contributions (ΔH ≪ 0)

Non-covalent interactions within the folded polypeptide release energy, yielding a highly favourable negative enthalpy change. These include:

  • Hydrogen Bonding: Internal hydrogen bonds form between the amide carbonyl oxygen and amide nitrogen of the polypeptide backbone (defining secondary structures), and between polar side chains.
  • Van der Waals Interactions: Highly packed hydrophobic cores optimize weak, short-range dispersion forces between non-polar side chains.
  • Electrostatic Interactions (Salt Bridges): Strong ionic attractions form between positively charged side chains (Lys, Arg, His) and negatively charged side chains (Asp, Glu) on the protein surface or interior.
Classic Biophysical Milestones and Paradoxes

2. Classic Biophysical Milestones and Paradoxes

Our understanding of how primary sequence determines tertiary structure, and the pathways through which folding occurs, is anchored in two foundational concepts.

2.1 Christian Anfinsen’s Ribonuclease A Experiment

In the early 1960s, Christian Anfinsen and his colleagues conducted pioneering experiments on bovine pancreatic Ribonuclease A (RNase A). RNase A is a stable, single-chain digestive enzyme containing 124 amino acid residues, with a molecular mass of 13,700 Da. Critically, its native, active structure is stabilized by four intramolecular disulfide linkages formed between eight specific cysteine residues.

Anfinsen's RNase A Refolding Pathways

Native RNase A (100% Active) Add 8 M Urea (Denaturant) & β-Mercaptoethanol (Reductant) - Destroys non-covalent bonds & cleaves disulfides Denatured, Reduced RNase A (Unfolded, 0% Active) Remove reductant FIRST, then remove urea LATER Remove BOTH denaturant and reductant SIMULTANEOUSLY Scrambled RNase A - Random disulfide pairing - ~1% Catalytic Activity Native RNase A - Correct disulfide pairing - 100% Catalytic Activity Add trace β-mercaptoethanol in the absence of urea (Catalyses disulfide shuffling)

Anfinsen's protocol demonstrated two distinct refolding pathways depending on the chemical microenvironment:

  • Denaturation and Reduction: Treatment of native RNase A with 8 M urea (which disrupts the non-covalent hydrogen bonding and hydrophobic networks) and β-mercaptoethanol (which reduces the four covalent disulfide bonds back to free sulfhydryl groups, -SH) completely unfolded the enzyme, reducing its catalytic activity to 0%.
  • Path A (Simultaneous Removal): When both urea and β-mercaptoethanol were removed simultaneously (or urea was dialysed out first, followed by the removal of the reductant), the polypeptide chain spontaneously folded back into its native, biologically active three-dimensional conformation. The eight cysteine residues paired correctly into their original four disulfide bonds. This proved that the primary amino acid sequence contains all the necessary information required to guide the folding of the polypeptide chain into its native three-dimensional structure.
  • Path B (Sequential Removal - Scrambling): When the reducing agent (β-mercaptoethanol) was removed first while the protein remained dissolved in 8 M urea, covalent disulfide bonds reformed in the presence of oxygen. However, because the denaturing urea prevented the polypeptide chain from adopting its native, low-energy conformation, the cysteine residues paired randomly. This yielded a heterogeneous mixture of inactive, misfolded molecules termed scrambled ribonuclease, which exhibited less than 1% of the wild-type activity.
  • Rescue of Scrambled Ribonuclease: Remarkably, when scrambled ribonuclease was exposed to trace amounts of β-mercaptoethanol in the absence of urea, the trace reductant acted as a catalyst to reversibly cleave the incorrect, thermodynamically unstable disulfide bonds. This allowed the polypeptide chain to continuously sample different conformations until it settled into the most thermodynamically stable, lowest-energy native state, which correctly reformed the native disulfide bonds.

The Mathematics of Random Disulfide Pairing in RNase A

If disulfide bond formation were a completely random, non-directed process, the probability of eight cysteine residues correctly forming four specific disulfide bonds can be calculated sequentially:

  • First Disulfide Bond: The first cysteine selected can pair with any of the remaining 7 cysteines. The probability of choosing the single correct partner is:
    P1 = 1 / 7
  • Second Disulfide Bond: With 6 cysteines remaining, the next chosen cysteine can pair with any of the remaining 5 residues. The probability of choosing the correct partner is:
    P2 = 1 / 5
  • Third Disulfide Bond: With 4 cysteines remaining, the next selected cysteine can pair with any of the remaining 3 residues. The probability of choosing the correct partner is:
    P3 = 1 / 3
  • Fourth Disulfide Bond: The final 2 cysteines must pair with each other. The probability of correct pairing is:
    P4 = 1 / 1 = 1

To find the total probability (Pcorrect) of randomly forming all four correct disulfide bonds simultaneously, we multiply the individual probabilities:

Pcorrect = 1/7 × 1/5 × 1/3 × 1 = 1/105 ≈ 0.0095 (0.95%)

This rigorous mathematical treatment matches the experimental observation that random, non-guided oxidation of reduced RNase A yields approximately 1% of biologically active enzyme, confirming that correct folding must guide the spatial alignment of the cysteines prior to covalent oxidation.

2.2 Levinthal’s Paradox and the Molten Globule State

In 1969, Cyrus Levinthal pointed out a striking paradox regarding protein folding kinetics. If a relatively small polypeptide chain of 100 amino acids were to find its native state by systematically sampling every possible conformation, the folding time would exceed the age of the universe.

Assuming each amino acid residue can adopt only three possible conformations (via rotations about the φ and ψ dihedral angles), a 100-residue peptide has 3100 ≈ 5 × 1047 potential conformations. If the polypeptide can transition between different conformations at the physical speed of bond rotation (10-13 seconds per conformation), the time required for an unbiased random search is:

Time = (5 × 1047 conformations) / (1013 conformations s−1)
= 5 × 1034 seconds ≈ 1.6 × 1027 years

Because proteins actually fold in vivo and in vitro within milliseconds to seconds, protein folding cannot be a random, trial-and-error search. Instead, it must progress through structured, thermodynamically directed folding pathways.

The Protein Folding Pathway

Unfolded Random Coil Fast Collapse (driven by Hydrophobic Effect) Molten Globule Intermediate • Highly compact • High native secondary structure content • Lacks locked tertiary packing • Highly mobile, liquid-like core Slow Reorganisation & Packing (Secondary optimization steps) Folded Native Protein

The key intermediate in these structured pathways is the molten globule state:

  • Hydrophobic Collapse: Folding is initiated by a spontaneous, incredibly fast collapse of hydrophobic residues into the interior of the protein to minimize solvent exposure, occurring within a few milliseconds.
  • Characteristics of the Molten Globule:
    • It is highly compact, closely resembling the volume of the native state.
    • It possesses a high content of native-like secondary structures (α-helices and β-sheets).
    • It lacks the specific, rigid side-chain packing characteristic of the fully folded tertiary structure.
    • The hydrophobic interior side chains remain highly mobile and disordered, behaving dynamically like a liquid rather than a solid.
The Molecular Chaperone Machinery

3. The Molecular Chaperone Machinery

Understanding the vital protein networks that actively guide folding, isolate intermediates, and dismantle aggregates to maintain cellular proteostasis.

3.1 Functional Taxonomy of Chaperones

While many proteins can fold spontaneously, the highly crowded intracellular environment (containing up to 300–400 mg/mL of macromolecules) dangerously exposes sticky, hydrophobic folding intermediates, directly leading to toxic, off-pathway protein aggregation. Cells proactively utilize highly conserved molecular chaperones to assist in folding, targeting, and actively maintaining proteostatic stability.

Molecular chaperones do not convey any new structural information to the folding protein, nor do they chemically alter the final thermodynamic stability of the native state. Instead, they prevent aggregation by selectively binding exposed hydrophobic surfaces, isolating fragile folding intermediates, or actively dissolving existing aggregates, rigorously utilizing the energy derived from ATP hydrolysis.

Chaperones are classified into three primary functional groups based precisely on their molecular mechanism of action:

Functional Classification of Molecular Chaperones

Molecular Chaperones Foldases - Actively refold proteins - ATP-dependent e.g., Hsp70, Hsp60 Holdases - Bind intermediates - Prevent aggregation e.g., sHsps, Hsp40 Disaggregases - Actively solubilise - ATP-dependent e.g., Hsp100, ClpB
  • Foldases: Chaperones that actively catalyze and strictly assist in the refolding of unfolded or partially folded polypeptide substrates directly using ATP-driven cycles of binding and release. Key biological examples include the Hsp70 and Hsp60 (chaperonins) families.
  • Holdases: Chaperones that selectively bind to exposed hydrophobic regions of unfolded proteins or delicate folding intermediates to securely keep them in a soluble state and actively prevent their toxic aggregation, but crucially do not actively refold them. Examples include small heat shock proteins (sHsps) and Hsp40.
  • Disaggregases: Highly specialized, ATP-dependent multimeric protein complexes that actively and forcefully extract individual polypeptide chains from massive pre-existing protein aggregates and feed them directly back into normal folding pathways. Key examples include the Hsp100 family in plants/yeast and bacterial ClpB.

Most chaperones are synthesized at normal basal levels under standard physiological conditions but are massively upregulated under severe cellular stress or heat shock (e.g., elevated temperatures, heavy metals), which physically denatures proteins. Thus, they are historically named Heat Shock Proteins (Hsps), conventionally categorized by their specific molecular weight (in kDa).

3.2 The Hsp70 Family Machinery

The Hsp70 family (including inducible Hsp70 and constitutively expressed Hsc70 in mammalian cells) acts very early in a protein's life cycle. They rapidly bind to extended, hydrophobic peptide segments of approximately seven amino acid residues the instant they exit the ribosome.

Structural Domains of Hsp70

  • N-terminal Nucleotide-Binding Domain (NBD): A ~40 kDa catalytic ATPase domain that actively hydrolyses ATP to ADP.
  • C-terminal Substrate-Binding Domain (SBD): A ~25 kDa target domain containing a hydrophobic binding pocket surrounded by a mobile helical "lid" structure. This domain preferentially binds hydrophobic peptide regions rich in Leu, Ile, Val, Phe, and Tyr.
  • Linker Region: A highly conserved, flexible hydrophobic linker that mechanically couples large conformational changes between the NBD and SBD.

The Hsp70 ATP-Regulated Binding Cycle

ATP-BOUND STATE (Low affinity, Open Lid) NBD-ATP SBD (Lid Open) Unfolded substrate binds/escapes rapidly Hsp40 (Co-chaperone) stimulates ATP hydrolysis ADP-BOUND STATE (High affinity, Closed Lid) NBD-ADP SBD (Lid Closed) Substrate locked in hydrophobic pocket, prevented from aggregating NEF (Nucleotide Exchange Factor) exchanges ADP for ATP NBD-ATP SBD (Lid Open) Substrate released to attempt folding

The Hsp70 Catalytic Cycle

The activity of Hsp70 is strictly regulated by a complex ATP-dependent conformational cycle:

  1. ATP-Bound State (Lid Open): When ATP is bound to the NBD, the SBD helical lid is physically open. Hydrophobic peptide substrates bind and dissociate with rapid kinetics; the absolute affinity for the substrate is extremely low.
  2. Hydrolysis (Lid Closed): The dedicated co-chaperone Hsp40 (DnaJ) selectively binds the unfolded substrate and directly presents it to Hsp70, simultaneously stimulating the intrinsic ATPase activity of the NBD. ATP is rapidly hydrolysed to ADP. This massive conformational change aggressively closes the helical lid of the SBD, trapping the substrate deep within the hydrophobic pocket with very high affinity.
  3. ADP-Bound State: The sensitive substrate is held securely, actively preventing it from interacting with other exposed hydrophobic proteins in the cytosol.
  4. Nucleotide Exchange and Release: A Nucleotide Exchange Factor (NEF) (such as GrpE in bacteria or BAG1 in eukaryotes) structurally binds the NBD, strongly prompting the release of ADP. A fresh molecule of ATP binds to the NBD. This causes the helical lid of the SBD to violently swing open, releasing the polypeptide. The released protein can either successfully fold into its native conformation, or, if hydrophobic regions remain exposed, re-enter the Hsp70 cycle or transfer to the Hsp60 chaperonin system.

3.3 The Hsp60 Family (Chaperonins)

The Hsp60 family of chaperonins physically forms massive, hollow, double-ringed complexes. They provide a secure physical isolation chamber (historically termed an "Anfinsen cage") that completely encapsulates a single folding polypeptide, strictly protecting it from the crowded cytosol and safely allowing it to fold without aggregating.

Phylogenetic Classification of Chaperonins

  • Group I Chaperonins: Found exclusively in bacteria (GroEL), mitochondria (Hsp60), and chloroplasts (Cpn60). They operate strictly in conjunction with a separate heptameric, dome-like co-chaperonin "cap" structure (GroES in bacteria, Hsp10/Cpn10 in organelles). The main chamber ring always consists of exactly seven identical subunits.
  • Group II Chaperonins: Found exclusively in archaea (thermosome) and the eukaryotic cytosol (TRiC/CCT). They consist of larger eight-membered rings and are entirely independent of separate co-chaperonins; they feature a built-in structural protrusion on each subunit that acts as an integrated lid, opening and closing in an ATP-dependent manner.

The GroEL/GroES Catalytic Cycle in Escherichia coli

GroEL is a massive 14-subunit homotetramer logically arranged as two back-to-back heptameric rings, functionally forming two independent biological chambers.

GroEL/ES Chaperonin Encapsulation

Unfolded Substrate Hydrophobic Chamber (GroEL Cis-ring) Inactive Chamber GroES Cap + ATP GroES Hydrophilic Chamber (Encapsulated) Inactive Chamber ATP Binding triggers internal charge shift, exposing polar lining

The full cycle of complex encapsulation and folding occurs strictly through the following steps:

  1. Substrate Binding: An unfolded protein critically exposes large hydrophobic patches that bind strongly to a hydrophobic ring lining the interior rim of the empty "cis" chamber of GroEL. At this specific point, the cis-ring is actively in a low-affinity state for the GroES cap.
  2. ATP and Cap Binding: Seven ATP molecules rapidly bind to the individual subunits of the cis-ring. This dynamically induces a massive conformational shift that strongly recruits the heptameric GroES cap, which rapidly binds to the rim and physically seals the chamber shut.
  3. Conformational Shift (Anfinsen Cage Activation): The structural binding of GroES and ATP physically causes a dramatic 60-degree mechanical rotation of the GroEL apical domains. This functionally doubles the internal volume of the chamber and physically hides the hydrophobic binding lining, exposing polar, highly hydrophilic residues instead. The hydrophobic protein is aggressively forced into the aqueous center of the sealed chamber, actively initiating forced folding.
  4. ATP Hydrolysis: Over an approximate span of 10 seconds, the seven ATP molecules securely inside the cis-ring are chemically hydrolysed to ADP. This uniquely acts as a biological molecular timer, during which the target protein physically folds in complete isolation.
  5. Trans-Ring Triggering: Concurrently, a new unfolded polypeptide and seven new ATP molecules firmly bind to the opposite "trans" ring of GroEL.
  6. Disassembly and Release: The binding of fresh ATP to the trans-ring violently triggers the mechanical release of the GroES cap, the seven ADPs, and the folded target protein exactly from the cis-ring. If the released protein is still partially unfolded, it will expose hydrophobic residues, directly allowing it to quickly bind to another chaperonin and continuously repeat the cycle until fully folded.
Protein Degradation and Pathologies

4. Pathological Protein Misfolding: Amyloid Fibrils

Understanding the transition of monomeric proteins into highly structured, toxic amyloid aggregates and their role in degenerative diseases.

4.1 Structural Architecture of Amyloids

When protein folding quality control systems fail, hydrophobic proteins can self-assemble into highly structured, insoluble proteinaceous aggregates known as amyloid fibrils. These deposits accumulate in tissues, causing progressive cell death and tissue degeneration, collectively referred to as amyloidoses.

Amyloid Fibril Assembly Pathway

Monomeric Protein - Rich in α-helices/coils Oligomeric Intermediate (Highly toxic soluble species) Protofilament - Parallel/antiparallel β-sheets Amyloid Fibril - Rigid, insoluble cross-β sheet structure

Despite originating from structurally diverse monomeric precursor proteins with distinct amino acid sequences, all amyloid fibrils share a common structural motif:

  • The Cross-β Sheet: Under the electron microscope, amyloids appear as unbranched, rigid, linear fibrils 7 to 10 nm in diameter. X-ray fiber diffraction reveals that these fibrils consist of β-sheets where the individual β-strands run perpendicular to the long axis of the fibril, while the hydrogen bonds linking these strands run parallel to the axis.
  • Proteolytic Resistance: This highly ordered, extensively hydrogen-bonded structure packs tightly, excluding water. Consequently, amyloids are highly resistant to proteolytic degradation, heat denaturation, and common detergents.
  • Nucleation-Dependent Polymerisation Kinetics: Amyloid formation occurs via a nucleated growth pathway:
    • Lag Phase: The rate-limiting step during which monomeric proteins slowly assemble to form a stable oligomeric nucleus.
    • Elongation Phase: Once a stable nucleus (seed) is established, monomers are rapidly recruited to the growing ends of the protofilaments, leading to exponential fibril assembly.

4.2 Representative Amyloidoses and Pathological Proteins

Disease Affected Organ/Tissue Pathological Protein/Peptide Core Structure
Alzheimer's Disease Brain (Cortex, Hippocampus) Amyloid-β (Aβ40 / Aβ42) peptide cleaved from APP; Tau protein Extracellular senile plaques (cross-β); intracellular neurofibrillary tangles
Transmissible Spongiform Encephalopathies (TSEs) Brain (Cerebrum, Cerebellum) Prion Protein (PrPSc, mutated from normal PrPC) Rich in β-sheet; highly infectious; induces conformational conversion of normal PrP
Type II Diabetes Mellitus Pancreas (Islets of Langerhans) Amylin (Islet Amyloid Polypeptide, IAPP) Co-secreted with insulin; damages insulin-producing pancreatic β-cells
Parkinson's Disease Brain (Substantia Nigra) α-Synuclein Intracellular inclusions termed Lewy Bodies

5. Ubiquitin-Mediated Protein Degradation

Exploring the ATP-dependent pathways, structural degrons, and enzymatic cascades responsible for clearing misfolded or obsolete proteins.

5.1 Ubiquitin Structure and Signals

Protein degradation is essential for clearing damaged, aged, or misfolded proteins, and for regulating key cellular pathways (such as the cell cycle). In eukaryotic cells, this is mediated by the Ubiquitin-Proteasome System (UPS), an ATP-dependent pathway that targets specific proteins for degradation.

Ubiquitin is a highly conserved polypeptide consisting of 76 amino acid residues (molecular weight ∼8.6 kDa) found in all eukaryotic cells.

  • Degradation Signals (Degrons): Proteins destined for degradation contain specific structural motifs or exposed residues, termed degrons, which are recognized by the UPS machinery.
  • Lysine Architecture: Ubiquitin contains seven lysine residues (K6, K11, K27, K29, K33, K48, and K63). The mode of linkage determines the cellular fate:
    • K48-linked polyubiquitin chains: A chain of four or more ubiquitins linked via Lys-48 is the canonical signal targeting proteins specifically to the 26S proteasome for degradation.
    • K63-linked polyubiquitin chains: Typically signal non-proteolytic pathways, such as DNA repair, endocytosis, and kinase activation.
    • Monoubiquitination: The attachment of a single ubiquitin molecule, regulating receptor internalisation and histone modification.

5.2 The Enzymatic Ubiquitination Cascade

The covalent attachment of ubiquitin to a target protein occurs via a sequential, three-step enzymatic cascade requiring ATP hydrolysis.

The Three-Step Ubiquitination Pathway

[ STEP 1: Activation by E1 ] Ubiquitin + ATP Ubiquitin-AMP E1-Cys-S∼Ub (Thioester intermediate) − Releases PPi [ STEP 2: Conjugation by E2 ] E1-Cys-S∼Ub + E2-Cys-SH E1-Cys-SH + E2-Cys-S∼Ub (Transthioesterification) [ STEP 3: Ligation by E3 ] E2-Cys-S∼Ub + Target (Lys) E3 Target-Lys-NH-CO-Ub (Isopeptide bond)
  1. Ubiquitin Activation (E1): In an ATP-dependent reaction, the carboxy-terminal glycine residue (Gly-76) of ubiquitin is adenylated, releasing pyrophosphate (PPi). The carboxyl group of Gly-76 is then forcefully transferred to the sulfhydryl (-SH) group of a highly conserved cysteine residue within the active site of the E1 enzyme, formally establishing a high-energy thioester bond (E1-S∼Ub).
  2. Ubiquitin Conjugation (E2): The activated ubiquitin is seamlessly transferred from E1 directly to the active-site cysteine of a specialized ubiquitin-conjugating enzyme (E2) via a transthioesterification reaction, actively forming a new E2-S∼Ub intermediate. E1 is then released.
  3. Ubiquitin Ligation (E3): A dedicated ubiquitin ligase (E3) actively recognizes the specific degron on the target protein and dynamically brings it into close structural proximity with the E2-S∼Ub complex. The E3 enzyme then catalyses the permanent transfer of ubiquitin from E2 directly to the ε-amino group of a lysine residue on the target protein, permanently establishing a stable, covalent isopeptide bond.

Major Classes of E3 Ubiquitin Ligases

  • HECT (Homologous to E6-AP Carboxyl Terminus) Family: These powerful ligases catalyze ubiquitination in a strict two-step reaction. They contain their own active-site cysteine that transiently accepts ubiquitin from E2 (temporarily forming a covalent E3-S∼Ub thioester intermediate) before physically transferring it to the target protein.
  • RING (Really Interesting New Gene) Family: These ligases do not form a covalent intermediate with ubiquitin. Instead, they act purely as structural scaffolds, simultaneously binding both the E2-S∼Ub complex and the target protein to perfectly align them, facilitating the direct transfer of ubiquitin from E2 straight to the target lysine. A key biological example is the SCF complex (consisting of Skp1, Cullin-1, Rbx1 [the RING domain], and a variable F-box protein that dictates absolute substrate specificity).
The 26S Proteasome Degradative Chamber

6. The 26S Proteasome Degradative Chamber

The 26S proteasome is a massive, cylindrical 2.5 MDa multicatalytic protease complex located in both the cytoplasm and nucleus of eukaryotic cells. It consists of one central 20S catalytic core particle capped at one or both ends by a 19S regulatory particle.

26S Proteasome Structure

26S Proteasome Overall Architecture 19S Regulatory Cap Recognises polyubiquitin, Cleaves Ub, Unfolds protein using AAA+ ATPases α ring (7 subunits) β ring (7 subunits) β ring (7 subunits) α ring (7 subunits) Structural gating, prevents entry Catalytic core (Thr active sites) 19S Regulatory Cap 20S Core Particle

6.1 The 20S Core Particle

The 20S core particle is a hollow cylinder formed by the stacking of four heptameric rings, exhibiting α7β7β7α7 stoichiometry:

  • The Outer α-Rings: Formed by seven structurally related but distinct α-subunits (α1 to α7). These subunits are catalytically inactive; they act as a strictly gated entryway, blocking unstructured proteins from entering the proteolytic chamber.
  • The Inner β-Rings: Formed by seven distinct β-subunits (β1 to β7). These subunits internally line the interior of the cylinder, hosting the active protease sites. The proteasome is an N-terminal nucleophile (Ntn) hydrolase that dynamically uses the hydroxyl group of an N-terminal threonine residue exactly as the catalytic nucleophile.

Catalytic Specificities

Three specific β-subunits perform distinctly targeted proteolytic cleavages within the core:

Subunit Activity Type Cleavage Specificity
β1 Subunit Caspase-like / Peptidyl-glutamyl peptide-hydrolysing Cleaves peptide bonds on the carboxyl side of acidic residues (Asp, Glu).
β2 Subunit Trypsin-like activity Cleaves peptide bonds on the carboxyl side of basic residues (Arg, Lys).
β5 Subunit Chymotrypsin-like activity Cleaves peptide bonds on the carboxyl side of hydrophobic aromatic residues (Phe, Tyr, Trp).

6.2 The 19S Regulatory Particle

The 19S cap acts as the master controller for entry into the 20S core. It contains exactly 19 individual protein subunits, highly organized into two dedicated sub-assemblies:

  • The Base (9 Subunits): Contains six distinct AAA+ family ATPases (Rpt1–Rpt6) arranged symmetrically in a hexameric ring. These specialized ATPases forcibly use energy derived directly from ATP hydrolysis to physically unfold the target protein and thread it through the opened α-ring directly into the 20S catalytic chamber.
  • The Lid (10 Subunits): Contains specialized receptors that rigorously recognize K48-linked polyubiquitin degradation chains, and a specialized isopeptidase (deubiquitinating enzyme, DUB) that precisely cleaves the polyubiquitin chain entirely from the substrate, allowing the free ubiquitin molecules to be biologically recycled.
The N-End Rule Pathway

7. The N-End Rule Pathway

The in vivo half-life of intracellular proteins ranges from a few seconds to several days. This selective stability is governed by the N-end rule pathway, a specialized branch of the ubiquitin-proteasome system where the half-life of a protein is determined by the identity of its amino-terminal amino acid residue. This N-terminal signal is formally termed an N-degron.

UNSTABLE N-TERMINUS (e.g., Arg, Lys, Phe) Recognised by N-recognins (E3) Rapid UPS Degradation STABLE N-TERMINUS (e.g., Ala, Gly, Val) Ignored by N-recognins Prolonged Intracellular Life

7.1 Structural Elements of the N-Degron

A biologically active N-degron strictly consists of two essential, complementary structural features:

  1. A highly destabilising N-terminal amino acid residue.
  2. A nearby, highly accessible internal lysine residue located within an unstructured, flexible segment of the target protein, which functionally serves as the direct covalent site for polyubiquitination.

7.2 Destabilising Residues in Saccharomyces cerevisiae

In the yeast model S. cerevisiae, the strict biochemical relationship between the precise N-terminal residue and overall protein stability is systematically categorized into distinct stabilizing and destabilizing classes:

N-terminal Residue (X-β-gal) In Vivo Half-Life in S. cerevisiae Classification
Met, Gly, Ala, Ser, Thr, Val, Cys, Pro > 20 hours (Up to > 30 hours) Highly Stabilising
Tyr, His 10 minutes (Tyr), > 5 hours (His) Moderately Destabilising
Ile, Asp, Glu 30 minutes Highly Destabilising
Lys, Arg 3 minutes (Lys), 2 minutes (Arg) Extremely Destabilising (Type 1)
Phe, Leu, Trp, Tyr 3 minutes Extremely Destabilising (Type 2)
Asn, Gln 3 minutes (Asn), 10 minutes (Gln) Destabilising (Requires modification)

7.3 Hierarchical Processing of N-Degrons

Eukaryotes efficiently categorize structurally destabilizing N-terminal residues into three hierarchical tiers based on whether they are directly recognized by E3 ligases (N-recognins) or require sequential enzymatic modification first:

Hierarchical Modification & Degradation Cascade

Tertiary Destabilising Asparagine (Asn, N) Glutamine (Gln, Q) N-terminal Deamidase Secondary Destabilising Aspartate (Asp, D) Glutamate (Glu, E) Oxidised Cys ATE1 (Arginyl-tRNA transferase) Primary Destabilising - Type 1 (Basic): Lys, Arg, His - Type 2 (Hydrophobic): Phe, Leu, Trp, Tyr, Ile Recognised by N-recognin E3s (e.g., UBR1 in yeast) Polyubiquitinated & Degraded (UPS)
  • Primary Destabilising Residues: Directly recognized and instantly targeted by dedicated E3 ubiquitin ligases (N-recognins, e.g., UBR1 in yeast).
    • Type 1 (Basic/Charged): Lysine (Lys), Arginine (Arg), and Histidine (His).
    • Type 2 (Bulky Hydrophobic): Phenylalanine (Phe), Leucine (Leu), Tryptophan (Trp), Tyrosine (Tyr), and Isoleucine (Ile).
  • Secondary Destabilising Residues: Aspartate (Asp) and Glutamate (Glu). To be efficiently recognized by the degradative machinery, they must first be enzymatically modified by arginyl-tRNA-protein transferase (ATE1), which selectively transfers a primary destabilizing arginine residue directly from Arg-tRNA onto the free N-terminus of the target protein.
  • Tertiary Destabilising Residues: Asparagine (Asn) and Glutamine (Gln). They must first undergo rapid enzymatic deamidation of their side-chain amide groups mediated by N-terminal deamidases. This converts them into the secondary destabilizing residues Aspartate and Glutamate, which are then subsequently arginylated by ATE1 and rapidly degraded.
Protein Sequencing: N-Terminal Analysis and Edman Degradation

8. Protein Sequencing: N-Terminal Analysis and Edman Degradation

Determining the primary amino acid sequence of a purified protein is a core biochemical technique, accomplished by selective N-terminal labelling or sequential Edman degradation.

8.1 Reagents for N-Terminal Identification

Before sequencing a polypeptide, identifying the very first N-terminal residue can critically confirm overall protein purity and biological identity. This is structurally achieved using two main electrophilic reagents:

Sanger’s Reagent (1-Fluoro-2,4-dinitrobenzene, FDNB)

Under mildly alkaline conditions (pH 9.5), the unprotonated N-terminal α-amino group attacks the aromatic carbon of FDNB via nucleophilic aromatic substitution, releasing hydrofluoric acid (HF) and forming a bright yellow dinitrophenyl (DNP)-peptide derivative.

Sanger's Reagent N-Terminal Labelling FDNB (1-Fluoro-2,4-dinitrobenzene) + H2N–CH(R1)–CO–Peptide Target Polypeptide Chain pH 9.5 − HF DNP-Peptide 6 M HCl Hydrolysis (Cleaves peptide bonds) DNP-Amino Acid (Intense Yellow Marker) + Free Amino Acids

When the stable DNP-peptide is subjected to robust strong acid hydrolysis (6 M HCl), virtually all internal peptide bonds are mechanically cleaved. However, the unique covalent carbon-nitrogen bond directly linking the dinitrophenyl group to the N-terminal amino acid remains completely intact. The intensely yellow DNP-amino acid is then selectively isolated and identified by chromatography.

Dansyl Chloride (5-Dimethylaminonaphthalene-1-sulfonyl chloride)

Operating via a mechanistically similar pathway, dansyl chloride actively reacts with the N-terminal amino group to form a stable dansyl-peptide intermediate. Following aggressive acid hydrolysis, the resulting dansyl-amino acid exhibits intense yellow-green fluorescence under ultraviolet (UV) light. This powerful optical property enables highly sensitive analytical detection at nanomolar concentrations, far exceeding the baseline sensitivity of standard Sanger's reagent.

8.2 The Edman Degradation Sequencing Chemistry

Originally developed by Pehr Edman, this brilliant chemical method sequentially and systematically removes exactly one amino acid residue at a time directly from the N-terminus of a peptide without inadvertently cleaving the remaining peptide bonds. It proceeds rigorously through a distinct three-stage reaction cycle:

The Edman Degradation 3-Stage Cycle [ STAGE 1: Coupling at pH 9 ] PITC + H2N–CH(R1)–CONH–Peptide (Edman Reagent) PTC-Peptide [ STAGE 2: Cyclisation / Cleavage ] PTC-Peptide Anhydrous TFA ATZ-Amino Acid (Extracted to organic layer) + Shortened Peptide (R2) Returns for next cycle of sequencing [ STAGE 3: Conversion ] ATZ-Amino Acid Aqueous H3O+ PTH-Amino Acid Stable derivative, identified via HPLC
  1. Stage 1: Coupling (Alkaline Phase): The uncharged N-terminal amino group of the peptide proactively reacts with phenylisothiocyanate (PITC, widely known as Edman's reagent) in a controlled alkaline solution (pH 9.0) to successfully form a stable phenylthiocarbamyl-peptide (PTC-peptide).
  2. Stage 2: Cyclisation and Cleavage (Anhydrous Acid Phase): The newly formed PTC-peptide is carefully treated with a strong anhydrous acid, such as trifluoroacetic acid (TFA). The sulfur atom of the phenylthiocarbamyl group biochemically attacks the carbonyl carbon of the first peptide bond, rapidly forming a cyclic intermediate. This selectively and exclusively cleaves the very first peptide bond, releasing the N-terminal residue strictly as an anilinothiazolinone-amino acid (ATZ-amino acid) derivative. Importantly, because specifically anhydrous acid is uniquely used, the remaining interior peptide bonds in the shortened polypeptide chain remain completely structurally intact.
  3. Stage 3: Conversion (Aqueous Acid Phase): The structurally unstable ATZ-amino acid is selectively and completely extracted into a non-polar organic solvent and chemically treated with dilute aqueous acid (H3O+). This strongly prompts an intramolecular rearrangement directly into a highly stable phenylthiohydantoin-amino acid (PTH-amino acid). The stable PTH-amino acid is then definitively identified by high-performance liquid chromatography (HPLC) or thin-layer chromatography, while the remaining intact, shortened polypeptide is subsequently subjected to another fresh round of Edman degradation.

9. Chemical and Enzymatic Cleavage Strategies

Standard Edman degradation can reliably sequence polypeptides of up to only 50 residues before signal loss occurs due to incomplete reactions. To sequence larger proteins, the polypeptide must first be cleaved into smaller peptides using sequence-specific endopeptidases or chemical reagents.

9.1 Enzymatic Cleavage Specificity (Proteolytic Enzymes)

Peptide Cleavage Specificity Map

- - - NH–CH(Rn-1)–CO NH–CH(Rn)–CO - - - (Amino Side) (Carboxyl Side)
Proteolytic Enzyme Class Cleavage Site Specificity Key Restrictions
Trypsin Endopeptidase Carboxyl side of basic residues: Lys, Arg (Rn-1 = Lys/Arg) Will not cleave if next residue is Proline (Rn = Pro)
Chymotrypsin Endopeptidase Carboxyl side of aromatic/large hydrophobic residues: Tyr, Phe, Trp (Rn-1 = Tyr/Phe/Trp) Will not cleave if next residue is Proline (Rn = Pro)
Pepsin Endopeptidase Amino side of aromatic/hydrophobic residues: Tyr, Phe, Trp, Leu (Rn = Tyr/Phe/Trp/Leu) Will not cleave if preceding residue is Proline (Rn-1 = Pro)
Elastase Endopeptidase Carboxyl side of small neutral residues: Ala, Gly, Ser (Rn-1 = Ala/Gly/Ser) Will not cleave if next residue is Proline (Rn = Pro)
Thermolysin Endopeptidase Amino side of bulky hydrophobic residues: Leu, Ile, Val, Phe (Rn = Leu/Ile/Val/Phe) Highly heat-stable metalloprotease
Carboxypeptidase A Exopeptidase Sequentially cleaves single residues from the C-terminus Will not cleave C-terminal Pro, Lys, Arg
Carboxypeptidase B Exopeptidase Sequentially cleaves single residues from the C-terminus Only cleaves C-terminal Arg, Lys
Carboxypeptidase C Exopeptidase Sequentially cleaves single residues from the C-terminus Cleaves any C-terminal residue
Aminopeptidase M Exopeptidase Sequentially cleaves single residues from the free N-terminus Cleaves all free N-terminal residues

9.2 Chemical Cleavage Specificity

  • Cyanogen Bromide (CNBr): Specifically cleaves peptide bonds strictly on the carboxyl side of methionine (Met) residues. The chemical reaction uniquely converts the C-terminal methionine residue of the newly released peptide into a cyclic peptidyl peptidyl-homoserine lactone.
  • Hydroxylamine (NH2OH): Specifically cleaves the peptide bonds linking asparagine and glycine (Asn-Gly) residues.

10. Biophysical Methods for Disulfide Bond Analysis

Many native extracellular proteins contain stabilizing disulfide bonds that rigidly link cysteine residues. Before accurate sequencing can occur, these covalent bonds must be chemically identified or permanently cleaved to allow the polypeptide chain to unfold completely.

Disulfide Bond Cleavage Strategies

Reduction/Alkylation vs. Oxidation Pathways

Protein–Cys–S–S–Cys–Protein Add DTT or β-mercaptoethanol (Reduces S-S to free -SH groups) Protein–Cys–SH + HS–Cys–Protein Add Iodoacetate (Irreversible alkylation) Protein–Cys–S–CH2–COO +      OOC–CH2–S–Cys–Protein (Stable Carboxymethyl-cysteines) Add Performic Acid (Strong Oxidation) 2 × Protein–Cys–SO3 (Stable Cysteic Acid residues, completely blocks S-S reform)
  • Performic Acid Oxidation: Treatment of a protein with strong performic acid actively oxidizes all disulfide bonds and free sulfhydryl groups completely into highly stable cysteic acid residues (-CH2-SO3). Because these new cysteic acid side chains carry a strong, permanent negative charge, intense electrostatic repulsion actively and permanently prevents the disulfide bonds from ever reforming.
  • Reduction and Alkylation: Alternatively, disulfide bonds are chemically reduced to free sulfhydryl groups using an excess of dithiothreitol (DTT) or β-mercaptoethanol. Because these newly reduced sulfhydryl groups are highly reactive and will spontaneously oxidize back into disulfides in the ambient presence of atmospheric oxygen, they must be irreversibly chemically blocked. This is elegantly achieved by adding iodoacetate, which covalently alkylates the free sulfhydryl groups to permanently form stable carboxymethyl-cysteine residues.

11. Analytical Protein Assays and Quantification Methods

Determining total protein concentration in an unknown sample is a fundamental, daily biochemical task, utilizing either direct UV spectroscopy or specialized dye-binding colorimetric assays.

11.1 Ultraviolet Spectroscopy (Direct Quantification)

Proteins inherently absorb ultraviolet light in the targeted range of 190 nm to 300 nm through two entirely distinct biophysical molecular mechanisms:

  • Peptide Backbone Absorption (190 nm – 210 nm): The peptide amide bond itself absorbs intensely in the far-UV spectrum (max at 205 nm). This assay is extremely sensitive, but is highly prone to severe baseline interference from many common buffer salts and laboratory solvents.
  • Aromatic Side Chain Absorption (280 nm): The aromatic amino acids tryptophan (Trp, W) and tyrosine (Tyr, Y) exhibit very strong intrinsic UV absorption in the near-UV spectrum, strongly peaking at exactly 280 nm. Tryptophan's intrinsic molar absorptivity at 280 nm is approximately four times greater than that of tyrosine. Phenylalanine strictly absorbs weakly at 257 nm and does not physically contribute significantly at the 280 nm wavelength.
  • The Nucleic Acid Correction: Contaminating nucleic acids absorb massively at 260 nm (due to their purine and pyrimidine rings), which strongly heavily skews direct protein measurements. To accurately calculate true protein concentration in mixed samples contaminated with nucleic acids, the raw absorbance at both 280 nm and 260 nm is accurately measured and mathematically corrected using the empirical Warburg-Christian equation:
    Protein Concentration (mg/mL) = 1.55 × A280 − 0.76 × A260

11.2 Classical Colorimetric Assays

When direct UV measurements are impractical due to buffer interference, sensitive visible-light colorimetric assays are used:

  • The Biuret Method: In a strongly alkaline solution, divalent copper ions (Cu2+) biochemically coordinate strictly with the nitrogen atoms of four adjacent peptide bonds, cleanly forming a distinct purple-coloured coordination complex. This assay is highly specific for proteins, but requires relatively high initial protein concentrations (low sensitivity).
  • The Lowry Method (Folin Assay): This sensitive method smoothly combines the initial Biuret copper-coordination reaction directly with the secondary reduction of the Folin-Ciocalteu reagent (phosphomolybdate-phosphotungstate). The active copper-treated peptide backbone, directly along with exposed tyrosine and tryptophan residues, powerfully reduces the Folin reagent to a deep blue heteropolymolybdenum complex. This assay is highly sensitive but remains heavily sensitive to false interference from reducing agents, standard detergents, and common laboratory buffers.
  • The Bradford Assay (Coomassie Dye-Binding): This incredibly rapid, highly sensitive modern assay completely relies on the direct structural binding of Coomassie Brilliant Blue G-250 dye to proteins, binding particularly tightly to basic (Lys, Arg, His) and aromatic residues. Upon physical binding, the dye instantly transitions from its protonated red/brown state (absorbing strongly at 465 nm) directly to its stable, unprotonated blue anionic state, completely shifting its optical absorption maximum to exactly 595 nm. The total measured change in absorbance at 595 nm is mathematically directly proportional to total protein concentration.
    Coomassie G-250 (Red-brown, 465 nm)Protein Binding ⟶ Coomassie G-250 (Blue, 595 nm)
  • The Ninhydrin Test (Amino Acid Identification): Ninhydrin (triketohydrindene hydrate) is a powerful chemical oxidizing agent that cleanly reacts exclusively with the free α-amino group of individual amino acids or short peptides. The reaction proceeds through active oxidative deamination, releasing free ammonia (NH3), carbon dioxide (CO2), and an aldehyde. The released ammonia then instantly condenses with one molecule of reduced ninhydrin (hydrindantin) and one molecule of oxidized ninhydrin to form a deep blue-purple complex historically known as Ruhemann's purple, which absorbs strongly at 570 nm.
    • The Proline Exception: Because proline is structurally a secondary cyclic amine (an imino acid), it physically cannot undergo oxidative deamination. Instead, it chemically condenses with ninhydrin to form a entirely distinct, bright yellow-coloured product, absorbing strictly at 440 nm.

12. Quantitative Protein Purification and Specific Activity

Isolating a single target protein from a highly complex cell lysate requires a precise series of purification steps based on size, charge, solubility, or binding affinity. To mathematically monitor the success of any purification protocol, both total protein concentration and the specific enzymatic activity of the target protein must be quantified accurately at each step.

12.1 Mathematical Formulas for Purification Analysis

  • Total Protein (mg): The total absolute mass of all combined protein in the isolated fraction.
    Total Protein (mg) = Protein Concentration (mg/mL) × Total Volume (mL)
  • Total Activity (Units or U): The total measurable catalytic capability of the target enzyme strictly within the fraction.
    Total Activity (U) = Enzyme Activity Concentration (U/mL) × Total Volume (mL)
  • Specific Activity (U/mg): A highly critical mathematical measure of active enzyme purity, accurately representing the absolute units of target enzyme activity per total milligram of total sample protein.
    Specific Activity = Total Activity (Units) / Total Protein (mg)
    As a protein is actively purified and inactive contaminating proteins are mechanically removed, the Specific Activity must actively increase, reaching an absolute maximum plateau when the target protein is 100% pure.
  • Yield (%): The absolute percentage of the starting target enzyme activity successfully recovered after a purification step.
    Yield (%) = [ Total Activity in Current Step (U) / Total Activity in Initial Crude Extract (U) ] × 100%
  • Purification Level (Fold): The measured numerical increase in specific activity (purity) precisely relative to the initial, unpurified crude extract.
    Purification Level (Fold) = Specific Activity in Current Step (U/mg) / Specific Activity in Initial Crude Extract (U/mg)

13. Fully-Worked Biochemical Analysis Problems

Exhaustive step-by-step solutions to complex biochemical peptide cleavage, sequencing, and quantitative purification problems.

Problem 1: Cleavage Site Selectivity of Trypsin

Question

Which peptide bond(s) marked as a, b, c, d, and e will be physically broken when the following oligopeptide is treated completely with trypsin at pH 7.0?

Lys a∼ Arg b∼ Pro c∼ Lys d∼ Arg e∼ Gly
Step-by-Step Resolution
  1. Recall Trypsin's Specificity: Trypsin is a highly specific endopeptidase that cleaves peptide bonds exclusively on the carboxyl side of basic amino acids, namely Lysine (Lys) and Arginine (Arg).
  2. Identify Potential Cleavage Sites:
    • Bond a: Located cleanly on the carboxyl side of Lys1. This is a potential target cleavage site.
    • Bond b: Located cleanly on the carboxyl side of Arg2. This is a potential target cleavage site.
    • Bond c: Located cleanly on the carboxyl side of Pro3. Proline is not basic; this bond absolutely cannot be cleaved by trypsin.
    • Bond d: Located cleanly on the carboxyl side of Lys4. This is a potential target cleavage site.
    • Bond e: Located cleanly on the carboxyl side of Arg5. This is a potential target cleavage site.
  3. Apply Steric Restrictions: Trypsin physically cannot cleave a Lys or Arg peptide bond if the immediately subsequent amino acid residue (Rn) is Proline.
    • For Bond a, the subsequent residue is Arg2. Cleavage is permitted.
    • For Bond b, the subsequent residue is Pro3. Cleavage is completely sterically blocked by the massive proline ring.
    • For Bond d, the subsequent residue is Arg5. Cleavage is permitted.
    • For Bond e, the subsequent residue is Gly6. Cleavage is permitted.
  4. Final Deduction: Trypsin will successfully cleave peptide bonds a, d, and e. (Note: If the starting peptide is cleaved sequentially, initial cleavage at 'a' successfully splits Lys1 directly from the peptide, leaving Arg2 at the new N-terminus, but because Pro3 remains adjacent, bond 'b' still remains completely intact).

Problem 2: Sequencing by Peptide Fragment Overlap

Question

A target polypeptide was exhaustively cleaved into two smaller peptides by Cyanogen Bromide (CNBr), and strictly into two different peptides by Trypsin. The isolated peptide fragments have the following sequences:

  • CNBr 1: Gly–Thr–Lys–Ala–Glu
  • CNBr 2: Ser–Met
  • Trypsin 1: Ser–Met–Gly–Thr–Lys
  • Trypsin 2: Ala–Glu

Determine the complete, contiguous sequence of the original parent peptide.

Step-by-Step Resolution
  1. Analyze Cleavage Specificities:
    • CNBr cleaves entirely at the carboxyl side of Methionine (Met). Therefore, any resulting peptide ending precisely in Met (or its homoserine lactone derivative) must definitively be an internal cleavage product, while the final peptide completely lacking Met must structurally represent the true C-terminus of the parent chain. Thus, CNBr 2 (Ser–Met) is positioned before CNBr 1 (Gly–Thr–Lys–Ala–Glu).
    • Trypsin cleanly cleaves on the carboxyl side of Lys and Arg. Therefore, Trypsin 1 (Ser–Met–Gly–Thr–Lys) must precede Trypsin 2 (Ala–Glu).
  2. Align the Overlapping Fragments: We logically arrange the fragmented sequences in an overlapping set to align the matching amino acid sequences:
    CNBr 2:        Ser–Met
    Trypsin 1:     Ser–Met–Gly–Thr–Lys
    CNBr 1:                Gly–Thr–Lys–Ala–Glu
    Trypsin 2:                         Ala–Glu
  3. Deduce Parent Sequence: By continuously reading the perfectly aligned consensus sequence directly from N-terminus to C-terminus, we obtain:
    Ser–Met–Gly–Thr–Lys–Ala–Glu

Problem 3: Enkephalin Peptide Net Charge and Digestion Analysis

Question

Enkephalins are naturally occurring short pentapeptides that actively act as endogenous opiates. Suppose a specific enkephalin has the following primary sequence:

Tyr–Gly–Gly–Phe–Ala–Met

(Note: While classical enkephalins are standard pentapeptides, this unique synthetic analogue is a hexapeptide containing an additional C-terminal Methionine).

Part a: What is the exact net charge of this peptide at pH 1.0 and pH 7.0?

Part b: Indicate the exact number of bonds that could be hydrolysed (if any) by complete digestion with: Trypsin, Chymotrypsin, and Cyanogen Bromide (CNBr).

Step-by-Step Resolution for Part a (Net Charge)
  1. Identify Ionisable Groups in Tyr–Gly–Gly–Phe–Ala–Met:
    • The N-terminal α-amino group of Tyrosine (pKa ≈ 9.0).
    • The C-terminal α-carboxyl group of Methionine (pKa ≈ 2.0).
    • The phenolic hydroxyl group on the specific side chain of Tyrosine (pKR ≈ 10.1).
    • All other internal residues (Gly, Gly, Phe, Ala, Met) have completely non-ionisable side chains.
  2. Calculate Net Charge at pH 1.0:
    At pH 1.0 (strongly acidic environment, pH < pKa of all chemical groups):
    • N-terminal α-amino group is fully protonated: -NH3+ (Charge = +1).
    • C-terminal α-carboxyl group is fully protonated: -COOH (Charge = 0).
    • Tyrosine side chain is fully protonated: -OH (Charge = 0).
    Net Charge at pH 1.0 = +1
  3. Calculate Net Charge at pH 7.0:
    At pH 7.0 (neutral environment, pKα-carboxyl < pH < pKα-amino, pKR):
    • N-terminal α-amino group remains fully protonated: -NH3+ (Charge = +1).
    • C-terminal α-carboxyl group is heavily deprotonated: -COO (Charge = -1).
    • Tyrosine side chain remains protonated: -OH (Charge = 0).
    Net Charge at pH 7.0 = 0 (Zwitterionic state)

Step-by-Step Resolution for Part b (Digestion)
  1. Digestion with Trypsin: Trypsin cleaves only on the carboxyl side of basic Lys and Arg. Since this entire hexapeptide contains absolutely neither Lysine nor Arginine, Trypsin will cleave 0 bonds.
  2. Digestion with Chymotrypsin: Chymotrypsin cleanly cleaves on the carboxyl side of bulky aromatic residues: Tyrosine (Tyr) and Phenylalanine (Phe).
    • Site 1: Carboxyl side of Tyr1. This specific peptide bond (Tyr–Gly) will be heavily cleaved.
    • Site 2: Carboxyl side of Phe4. This specific peptide bond (Phe–Ala) will be heavily cleaved.
    Therefore, Chymotrypsin will cleave exactly 2 bonds, cleanly yielding three separate fragments: Tyr, Gly–Gly–Phe, and Ala–Met.
  3. Digestion with Cyanogen Bromide (CNBr): CNBr specifically cleaves strictly on the carboxyl side of Methionine (Met) residues.
    • In this sequence, Met6 is the absolute C-terminal amino acid.
    • Because Met6 sits exactly at the C-terminus, there is completely no adjacent peptide bond extending from its carboxyl group left to cleave.
    Therefore, CNBr will cleave 0 bonds.

Problem 4: Specific Activity and Purification Yield Analysis

Question

Calculate the mathematically required Specific Activity and Purification Level (Fold) at each progressive step of the purification protocol detailed in the raw data table below:

Step Purification Procedure Total Protein (mg) Total Activity (Units)
1Crude Cell Extract20,0004,000,000
2Salt Precipitation5,0003,000,000
3DEAE-Cellulose Chromatography1,5001,000,000
4Size-Exclusion Chromatography500750,000
5Affinity Chromatography45675,000
Step-by-Step Calculations

Step 1: Crude Cell Extract

  • Specific Activity = 4,000,000 U / 20,000 mg = 200 U/mg
  • Purification Level = 200 U/mg / 200 U/mg = 1.0-fold (Starting Reference)
  • Yield = (4,000,000 U / 4,000,000 U) × 100% = 100%

Step 2: Salt Precipitation

  • Specific Activity = 3,000,000 U / 5,000 mg = 600 U/mg
  • Purification Level = 600 U/mg / 200 U/mg = 3.0-fold
  • Yield = (3,000,000 U / 4,000,000 U) × 100% = 75%

Step 3: DEAE-Cellulose Chromatography (Ion Exchange)

  • Specific Activity = 1,000,000 U / 1,500 mg ≈ 666.67 U/mg
  • Purification Level = 666.67 U/mg / 200 U/mg ≈ 3.33-fold
  • Yield = (1,000,000 U / 4,000,000 U) × 100% = 25%

Step 4: Size-Exclusion Chromatography (Gel Filtration)

  • Specific Activity = 750,000 U / 500 mg = 1,500 U/mg
  • Purification Level = 1,500 U/mg / 200 U/mg = 7.5-fold
  • Yield = (750,000 U / 4,000,000 U) × 100% = 18.75%

Step 5: Affinity Chromatography

  • Specific Activity = 675,000 U / 45 mg = 15,000 U/mg
  • Purification Level = 15,000 U/mg / 200 U/mg = 75.0-fold
  • Yield = (675,000 U / 4,000,000 U) × 100% = 16.875%

Compiled Purification Summary Table

Step Procedure Total Protein (mg) Total Activity (U) Specific Activity (U/mg) Yield (%) Purification (Fold)
1Crude Extract20,0004,000,000200100%1.0 (Ref)
2Salt PPT5,0003,000,00060075.0%3.0
3DEAE-Cellulose1,5001,000,00066725.0%3.3
4Size-Exclusion500750,0001,50018.8%7.5
5Affinity Chrom.45675,00015,00016.9%75.0

Conclusion: The final affinity chromatography step provides the most massive mathematical increase in purity, bringing the active enzyme to a remarkable 75-fold purity relative to the starting material, while securely retaining exactly 16.9% of the initial starting catalytic activity.

In this topic

Scroll to Top