Protein Folding
1. The Physics and Thermodynamics of Protein Folding
Protein folding is the highly coordinated physical process by which a linear, unstructured polypeptide chain folds into its unique, biologically active, and stable three-dimensional conformation, termed the native state.
[ Unstructured Polypeptide ] ====================> [ Native Conformation ]
High conformational entropy Low conformational entropy
High free energy (G) Minimum free energy (G)
Thermodynamically unstable Thermodynamically stable
The native conformation is dictated entirely by the primary amino acid sequence. Failure of a polypeptide to adopt its native state results in non-functional, inactive, or toxic misfolded species. Although protein folding occurs spontaneously in vivo and in vitro for many proteins, the intracellular environment presents extreme macromolecular crowding, requiring molecular chaperones to prevent off-pathway aggregation.
1.1 Thermodynamic Drivers of Protein Folding
The transition of an unfolded polypeptide to a folded macromolecule is thermodynamically spontaneous, meaning the change in Gibbs free energy ($\Delta G_{\text{folding}}$) is negative:
$$ \Delta G_{\text{folding}} = \Delta H_{\text{folding}} – T\Delta S_{\text{folding}} < 0 $$
At first glance, the folding process appears to violate the Second Law of Thermodynamics, as organizing a highly flexible, random-coil polymer into a single, well-defined conformation represents a massive decrease in the conformational entropy of the polypeptide chain ($\Delta S_{\text{polypeptide}} \ll 0$). This unfavourable entropic penalty is overcome by two major thermodynamic factors:
- The Hydrophobic Effect (Entropic Driver): This is the predominant thermodynamic force driving protein folding. In an unfolded state, hydrophobic side chains (such as Leu, Ile, Val, Phe, Trp) are exposed to the bulk aqueous solvent. This exposure forces surrounding water molecules to organize into highly structured, rigid, cage-like structures (clathrates) to maximize hydrogen bonding. This represents a significant local decrease in solvent entropy. Upon spontaneous collapse of the hydrophobic residues into the interior core of the protein, these ordered water molecules are released back into the bulk solvent, causing a massive increase in solvent entropy ($\Delta S_{\text{solvent}} \gg 0$). The net entropy change of the system (polypeptide + solvent) is highly positive:
$$ \Delta S_{\text{total}} = \Delta S_{\text{polypeptide}} + \Delta S_{\text{solvent}} > 0 $$ - Enthalpic Contributions ($\Delta H \ll 0$): Non-covalent interactions within the folded polypeptide release energy, yielding a highly favourable negative enthalpy change. These include:
- Hydrogen Bonding: Internal hydrogen bonds form between the amide carbonyl oxygen and amide nitrogen of the polypeptide backbone (defining secondary structures), and between polar side chains.
- Van der Waals Interactions: Highly packed hydrophobic cores optimize weak, short-range dispersion forces between non-polar side chains.
- Electrostatic Interactions (Salt Bridges): Strong ionic attractions form between positively charged side chains (Lys, Arg, His) and negatively charged side chains (Asp, Glu) on the protein surface or interior.
2. Classic Biophysical Milestones and Paradoxes
Our understanding of how primary sequence determines tertiary structure, and the pathways through which folding occurs, is anchored in two foundational concepts.
2.1 Christian Anfinsen’s Ribonuclease A Experiment
In the early 1960s, Christian Anfinsen and his colleagues conducted pioneering experiments on bovine pancreatic Ribonuclease A (RNase A). RNase A is a stable, single-chain digestive enzyme containing 124 amino acid residues, with a molecular mass of 13,700 Da. Critically, its native, active structure is stabilized by four intramolecular disulfide linkages formed between eight specific cysteine residues.
[ Native RNase A ]
(100% Active)
|
+---> Add 8 M Urea (Denaturant) & β-Mercaptoethanol (Reductant)
| - Destroys non-covalent interactions and cleaves disulfide bonds
v
[ Denatured, Reduced RNase A ]
(Unfolded, 0% Active)
|
+-----------------------+-----------------------+
| |
| Remove reductant FIRST, | Remove BOTH denaturant
| then remove urea LATER | and reductant SIMULTANEOUSLY
v v
[ Scrambled RNase A ] [ Native RNase A ]
- Random disulfide pairing - Correct disulfide pairing
- ~1% Catalytic Activity - 100% Catalytic Activity
| ^
+---> Add trace β-mercaptoethanol |
in the absence of urea -------------------+
- Catalyses disulfide shuffling to native state
Anfinsen’s protocol demonstrated two distinct refolding pathways depending on the chemical microenvironment:
- Denaturation and Reduction: Treatment of native RNase A with 8 M urea (which disrupts the non-covalent hydrogen bonding and hydrophobic networks) and $\beta$-mercaptoethanol (which reduces the four covalent disulfide bonds back to free sulfhydryl groups, $-\text{SH}$) completely unfolded the enzyme, reducing its catalytic activity to 0%.
- Path A (Simultaneous Removal): When both urea and $\beta$-mercaptoethanol were removed simultaneously (or urea was dialysed out first, followed by the removal of the reductant), the polypeptide chain spontaneously folded back into its native, biologically active three-dimensional conformation. The eight cysteine residues paired correctly into their original four disulfide bonds. This proved that the primary amino acid sequence contains all the necessary information required to guide the folding of the polypeptide chain into its native three-dimensional structure.
- Path B (Sequential Removal – Scrambling): When the reducing agent ($\beta$-mercaptoethanol) was removed first while the protein remained dissolved in 8 M urea, covalent disulfide bonds reformed in the presence of oxygen. However, because the denaturing urea prevented the polypeptide chain from adopting its native, low-energy conformation, the cysteine residues paired randomly. This yielded a heterogeneous mixture of inactive, misfolded molecules termed scrambled ribonuclease, which exhibited less than 1% of the wild-type activity.
- Rescue of Scrambled Ribonuclease: Remarkably, when scrambled ribonuclease was exposed to trace amounts of $\beta$-mercaptoethanol in the absence of urea, the trace reductant acted as a catalyst to reversibly cleave the incorrect, thermodynamically unstable disulfide bonds. This allowed the polypeptide chain to continuously sample different conformations until it settled into the most thermodynamically stable, lowest-energy native state, which correctly reformed the native disulfide bonds.
The Mathematics of Random Disulfide Pairing in RNase A
If disulfide bond formation were a completely random, non-directed process, the probability of eight cysteine residues correctly forming four specific disulfide bonds can be calculated sequentially:
- First Disulfide Bond: The first cysteine selected can pair with any of the remaining 7 cysteines. The probability of choosing the single correct partner is:
$$ P_1 = \frac{1}{7} $$ - Second Disulfide Bond: With 6 cysteines remaining, the next chosen cysteine can pair with any of the remaining 5 residues. The probability of choosing the correct partner is:
$$ P_2 = \frac{1}{5} $$ - Third Disulfide Bond: With 4 cysteines remaining, the next selected cysteine can pair with any of the remaining 3 residues. The probability of choosing the correct partner is:
$$ P_3 = \frac{1}{3} $$ - Fourth Disulfide Bond: The final 2 cysteines must pair with each other. The probability of correct pairing is:
$$ P_4 = \frac{1}{1} = 1 $$
To find the total probability ($P_{\text{correct}}$) of randomly forming all four correct disulfide bonds simultaneously, we multiply the individual probabilities:
$$ P_{\text{correct}} = \frac{1}{7} \times \frac{1}{5} \times \frac{1}{3} \times 1 = \frac{1}{105} \approx 0.0095 \ (0.95\%) $$
This rigorous mathematical treatment matches the experimental observation that random, non-guided oxidation of reduced RNase A yields approximately 1% of biologically active enzyme, confirming that correct folding must guide the spatial alignment of the cysteines prior to covalent oxidation.
2.2 Levinthal’s Paradox and the Molten Globule State
In 1969, Cyrus Levinthal pointed out a striking paradox regarding protein folding kinetics. If a relatively small polypeptide chain of 100 amino acids were to find its native state by systematically sampling every possible conformation, the folding time would exceed the age of the universe.
Assuming each amino acid residue can adopt only three possible conformations (via rotations about the $\phi$ and $\psi$ dihedral angles), a 100-residue peptide has $3^{100} \approx 5 \times 10^{47}$ potential conformations. If the polypeptide can transition between different conformations at the physical speed of bond rotation ($10^{-13}$ seconds per conformation), the time required for an unbiased random search is:
$$ \text{Time} = \frac{5 \times 10^{47} \text{ conformations}}{10^{13} \text{ conformations s}^{-1}} = 5 \times 10^{34} \text{ seconds} \approx 1.6 \times 10^{27} \text{ years} $$
Because proteins actually fold in vivo and in vitro within milliseconds to seconds, protein folding cannot be a random, trial-and-error search. Instead, it must progress through structured, thermodynamically directed folding pathways.
[ Unfolded Random Coil ]
|
| Fast Collapse (driven by Hydrophobic Effect)
v
[ Molten Globule Intermediate ]
- Highly compact
- High content of native secondary structure (α-helices, β-sheets)
- Lacks locked tertiary side-chain packing
- Highly mobile, liquid-like interior core
|
| Slow Reorganisation and Packing (Secondary steps)
v
[ Folded Native Protein ]
The key intermediate in these pathways is the molten globule state:
- Hydrophobic Collapse: Folding is initiated by a spontaneous, incredibly fast collapse of hydrophobic residues into the interior of the protein to minimize solvent exposure, occurring within a few milliseconds.
- Characteristics of the Molten Globule:
- It is highly compact, closely resembling the volume of the native state.
- It possesses a high content of native-like secondary structures ($\alpha$-helices and $\beta$-sheets).
- It lacks the specific, rigid side-chain packing characteristic of the fully folded tertiary structure.
- The hydrophobic interior side chains remain highly mobile and disordered, behaving like a liquid rather than a solid.
3. The Molecular Chaperone Machinery
While many proteins can fold spontaneously, the highly crowded intracellular environment (containing up to 300–400 mg/mL of macromolecules) exposes sticky, hydrophobic folding intermediates, leading to off-pathway aggregation. Cells use highly conserved molecular chaperones to assist in folding, targeting, and maintaining proteostatic stability.
Molecular chaperones do not convey new structural information to the folding protein, nor do they alter the thermodynamic stability of the native state. Instead, they prevent aggregation by binding hydrophobic surfaces, isolating folding intermediates, or actively dissolving aggregates, utilizing the energy of ATP hydrolysis.
3.1 Functional Taxonomy of Chaperones
Chaperones are classified into three primary functional groups based on their molecular mechanism of action:
Molecular Chaperones
|
+---------------------------+---------------------------+
| | |
[ Foldases ] [ Holdases ] [ Disaggregases ]
- Actively refold proteins - Bind intermediates - Actively solubilise
- ATP-dependent - Prevent aggregation - ATP-dependent
- e.g. Hsp70, Hsp60 - e.g. sHsps, Hsp40 - e.g. Hsp100, ClpB
- Foldases: Chaperones that actively catalyze and assist in the refolding of unfolded or partially folded polypeptide substrates using ATP-driven cycles of binding and release. Key examples include the Hsp70 and Hsp60 (chaperonins) families.
- Holdases: Chaperones that bind to exposed hydrophobic regions of unfolded proteins or folding intermediates to keep them in a soluble state and prevent their aggregation, but do not actively refold them. Examples include small heat shock proteins (sHsps) and Hsp40.
- Disaggregases: Specialized, ATP-dependent multimeric protein complexes that actively extract individual polypeptide chains from pre-existing protein aggregates and feed them back into folding pathways. Key examples include the Hsp100 family in plants/yeast and bacterial ClpB.
Most chaperones are synthesized at basal levels under normal physiological conditions but are significantly upregulated under heat shock or cellular stress (e.g., elevated temperatures, heavy metals), which denatures proteins. Thus, they are historically named Heat Shock Proteins (Hsps), categorized by their molecular weight (in kDa).
3.2 The Hsp70 Family Machinery
The Hsp70 family (including inducible Hsp70 and constitutively expressed Hsc70 in mammalian cells) acts early in a protein’s life cycle. They bind to extended, hydrophobic peptide segments of approximately seven amino acid residues as they exit the ribosome.
Structural Domains of Hsp70
- N-terminal Nucleotide-Binding Domain (NBD): A ~40 kDa ATPase domain that hydrolyses ATP to ADP.
- C-terminal Substrate-Binding Domain (SBD): A ~25 kDa domain containing a hydrophobic binding pocket surrounded by a helical “lid” structure. This domain preferentially binds hydrophobic peptide regions rich in Leu, Ile, Val, Phe, and Tyr.
- Linker Region: A highly conserved, flexible hydrophobic linker that couples conformational changes between the NBD and SBD.
ATP-BOUND STATE (Low affinity, Open Lid)
[ NBD-ATP ]========[ SBD (Lid Open) ] ---> Unfolded substrate binds/escapes rapidly
| Hsp40 (Co-chaperone) stimulates ATP hydrolysis
v
ADP-BOUND STATE (High affinity, Closed Lid)
[ NBD-ADP ]========[ SBD (Lid Closed) ] ---> Substrate locked in hydrophobic pocket,
prevented from aggregating
|
| NEF (Nucleotide Exchange Factor, e.g. GrpE) exchanges ADP for ATP
v
[ NBD-ATP ]========[ SBD (Lid Open) ] ---> Substrate released to attempt folding
The Hsp70 Catalytic Cycle
The activity of Hsp70 is regulated by an ATP-dependent conformational cycle:
- ATP-Bound State (Lid Open): When ATP is bound to the NBD, the SBD helical lid is open. Hydrophobic peptide substrates bind and dissociate with rapid kinetics; the affinity for substrate is low.
- Hydrolysis (Lid Closed): The co-chaperone Hsp40 (DnaJ) binds the substrate and presents it to Hsp70, simultaneously stimulating the intrinsic ATPase activity of the NBD. ATP is hydrolysed to ADP. This conformational change closes the helical lid of the SBD, trapping the substrate within the hydrophobic pocket with high affinity.
- ADP-Bound State: The substrate is held securely, preventing it from interacting with other exposed hydrophobic proteins in the cytosol.
- Nucleotide Exchange and Release: A Nucleotide Exchange Factor (NEF) (such as GrpE in bacteria or BAG1 in eukaryotes) binds the NBD, prompting the release of ADP. A new molecule of ATP binds to the NBD. This causes the helical lid of the SBD to swing open, releasing the polypeptide. The released protein can either fold into its native conformation, or, if hydrophobic regions remain exposed, re-enter the Hsp70 cycle or transfer to the Hsp60 chaperonin system.
3.3 The Hsp60 Family (Chaperonins)
The Hsp60 family of chaperonins forms large, hollow, double-ringed complexes. They provide a physical isolation chamber (termed an “Anfinsen cage”) that encapsulates a single folding polypeptide, protecting it from the crowded cytosol and allowing it to fold without aggregating.
Phylogenetic Classification of Chaperonins
- Group I Chaperonins: Found in bacteria (GroEL), mitochondria (Hsp60), and chloroplasts (Cpn60). They operate in conjunction with a heptameric, dome-like co-chaperonin “cap” (GroES in bacteria, Hsp10/Cpn10 in organelles). The chamber ring consists of seven identical subunits.
- Group II Chaperonins: Found in archaea (thermosome) and the eukaryotic cytosol (TRiC/CCT). They consist of eight-membered rings and are independent of co-chaperonins; they feature a built-in protrusion on each subunit that acts as a lid, opening and closing in an ATP-dependent manner.
The GroEL/GroES Catalytic Cycle in Escherichia coli
GroEL is a massive 14-subunit homotetramer arranged as two back-to-back heptameric rings, forming two independent chambers.
Unfolded Substrate GroES Cap
| |
v v
+-----------+ +===========+
| Hydro | | GroES |
| phobic | +-----------+
| Chamber | ===> | Hydro- | (ATP Hydrolysis triggers internal
| (GroEL) | | philic | charge shift, exposing polar lining)
+-----------+ | Chamber |
| Inactive | +-----------+
| Chamber | | Inactive |
+-----------+ +-----------+
Cis-ring Cis-ring (Encapsulated)
The cycle of encapsulation and folding occurs through the following steps:
- Substrate Binding: An unfolded protein exposes hydrophobic patches that bind to a hydrophobic ring lining the interior rim of the empty “cis” chamber of GroEL. At this point, the cis-ring is in a low-affinity state for the GroES cap.
- ATP and Cap Binding: Seven ATP molecules bind to the subunits of the cis-ring. This induces a conformational shift that recruits the heptameric GroES cap, which binds to the rim and seals the chamber.
- Conformational Shift (Anfinsen Cage Activation): The binding of GroES and ATP causes a dramatic rotation of the GroEL apical domains. This doubles the volume of the internal chamber and hides the hydrophobic binding lining, exposing polar, hydrophilic residues instead. The hydrophobic protein is forced into the aqueous center of the chamber, initiating folding.
- ATP Hydrolysis: Over approximately 10 seconds, the seven ATP molecules in the cis-ring are hydrolysed to ADP. This acts as a molecular timer, during which the protein folds in isolation.
- Trans-Ring Triggering: Concurrently, an unfolded polypeptide and seven new ATP molecules bind to the opposite “trans” ring of GroEL.
- Disassembly and Release: The binding of ATP to the trans-ring triggers the release of the GroES cap, the seven ADPs, and the folded protein from the cis-ring. If the released protein is still partially unfolded, it exposes hydrophobic residues, allowing it to quickly bind to another chaperonin and repeat the cycle.
4. Pathological Protein Misfolding: Amyloid Fibrils
When protein folding quality control systems fail, hydrophobic proteins can self-assemble into highly structured, insoluble proteinaceous aggregates known as amyloid fibrils. These deposits accumulate in tissues, causing progressive cell death and tissue degeneration, collectively referred to as amyloidoses.
Monomeric Protein ===> Oligomeric Intermediate ===> Protofilament ===> Amyloid Fibril
- Rich in α- (Highly toxic soluble - Parallel/ - Rigid, insoluble
helices/coils species) antiparallel cross-β sheet
β-sheets structure
4.1 Structural Architecture of Amyloids
Despite originating from structurally diverse monomeric precursor proteins with distinct amino acid sequences, all amyloid fibrils share a common structural motif:
- The Cross-$\beta$ Sheet: Under the electron microscope, amyloids appear as unbranched, rigid, linear fibrils 7 to 10 nm in diameter. X-ray fiber diffraction reveals that these fibrils consist of $\beta$-sheets where the individual $\beta$-strands run perpendicular to the long axis of the fibril, while the hydrogen bonds linking these strands run parallel to the axis.
- Proteolytic Resistance: This highly ordered, extensively hydrogen-bonded structure packs tightly, excluding water. Consequently, amyloids are highly resistant to proteolytic degradation, heat denaturation, and common detergents.
- Nucleation-Dependent Polymerisation Kinetics: Amyloid formation occurs via a nucleated growth pathway:
- Lag Phase: The rate-limiting step during which monomeric proteins slowly assemble to form a stable oligomeric nucleus.
- Elongation Phase: Once a stable nucleus (seed) is established, monomers are rapidly recruited to the growing ends of the protofilaments, leading to exponential fibril assembly.
4.2 Representative Amyloidoses and Pathological Proteins
| Disease | Affected Organ/Tissue | Pathological Protein/Peptide | Core Structure |
|---|---|---|---|
| Alzheimer’s Disease | Brain (Cortex, Hippocampus) | Amyloid-$\beta$ ($A\beta_{40}/A\beta_{42}$) peptide cleaved from APP; Tau protein | Extracellular senile plaques (cross-$\beta$); intracellular neurofibrillary tangles |
| Transmissible Spongiform Encephalopathies (TSEs) | Brain (Cerebrum, Cerebellum) | Prion Protein ($\text{PrP}^{\text{Sc}}$, mutated from $\text{PrP}^{\text{C}}$) | Rich in $\beta$-sheet; highly infectious; induces conformational conversion of normal PrP |
| Type II Diabetes Mellitus | Pancreas (Islets of Langerhans) | Amylin (Islet Amyloid Polypeptide, IAPP) | Co-secreted with insulin; damages insulin-producing pancreatic $\beta$-cells |
| Parkinson’s Disease | Brain (Substantia Nigra) | $\alpha$-Synuclein | Intracellular inclusions termed Lewy Bodies |
5. Ubiquitin-Mediated Protein Degradation
Protein degradation is essential for clearing damaged, aged, or misfolded proteins, and for regulating key cellular pathways (such as the cell cycle). In eukaryotic cells, this is mediated by the Ubiquitin-Proteasome System (UPS), an ATP-dependent pathway that targets specific proteins for degradation.
5.1 Ubiquitin structure
Ubiquitin is a highly conserved polypeptide consisting of 76 amino acid residues (molecular weight ~8.6 kDa) found in all eukaryotic cells.
- Degradation Signals (Degrons): Proteins destined for degradation contain specific structural motifs or exposed residues, termed degrons, which are recognized by the UPS machinery.
- Lysine Architecture: Ubiquitin contains seven lysine residues (K6, K11, K27, K29, K33, K48, and K63). The mode of linkage determines the cellular fate:
- K48-linked polyubiquitin chains: A chain of four or more ubiquitins linked via Lys-48 is the canonical signal targeting proteins to the 26S proteasome for degradation.
- K63-linked polyubiquitin chains: Typically signal non-proteolytic pathways, such as DNA repair, endocytosis, and kinase activation.
- Monoubiquitination: The attachment of a single ubiquitin molecule, regulating receptor internalisation and histone modification.
5.2 The Enzymatic Ubiquitination Cascade
The covalent attachment of ubiquitin to a target protein occurs via a sequential, three-step enzymatic cascade requiring ATP hydrolysis:
[ STEP 1: Activation by E1 ]
Ubiquitin + ATP =======> Ubiquitin-AMP =======> E1-Cys-S~Ub (Thioester intermediate)
- Releases PPi
[ STEP 2: Conjugation by E2 ]
E1-Cys-S~Ub + E2-Cys-SH =======> E1-Cys-SH + E2-Cys-S~Ub (Transthioesterification)
[ STEP 3: Ligation by E3 ]
E2-Cys-S~Ub + Target Protein (Lysine) === E3 ===> Target-Lys-NH-CO-Ub (Isopeptide bond)
- Ubiquitin Activation (E1): In an ATP-dependent reaction, the carboxy-terminal glycine residue (Gly-76) of ubiquitin is adenylated, releasing pyrophosphate ($\text{PP}_{\text{i}}$). The carboxyl group of Gly-76 is then transferred to the sulfhydryl ($-\text{SH}$) group of a conserved cysteine residue in the active site of E1, forming a high-energy thioester bond ($\text{E1-S}\sim\text{Ub}$).
- Ubiquitin Conjugation (E2): The activated ubiquitin is transferred from E1 to the active-site cysteine of a ubiquitin-conjugating enzyme (E2) via a transthioesterification reaction, forming an $\text{E2-S}\sim\text{Ub}$ intermediate. E1 is then released.
- Ubiquitin Ligation (E3): A ubiquitin ligase (E3) recognizes the degron on the target protein and brings it into proximity with the $\text{E2-S}\sim\text{Ub}$ complex. E3 catalyses the transfer of ubiquitin from E2 to the $\epsilon$-amino group of a lysine residue on the target protein, establishing a stable, covalent isopeptide bond.
Major Classes of E3 Ubiquitin Ligases
- HECT (Homologous to E6-AP Carboxyl Terminus) Family: These ligases catalyze ubiquitination in a two-step reaction. They contain an active-site cysteine that transiently accepts ubiquitin from E2 (forming a covalent $\text{E3-S}\sim\text{Ub}$ thioester intermediate) before transferring it to the target protein.
- RING (Really Interesting New Gene) Family: These ligases do not form a covalent intermediate with ubiquitin. Instead, they act as structural scaffolds, binding both the $\text{E2-S}\sim\text{Ub}$ complex and the target protein to align them, facilitating direct transfer of ubiquitin from E2 to the target lysine. A key example is the SCF complex (consisting of Skp1, Cullin-1, Rbx1 [the RING domain], and a variable F-box protein that dictates substrate specificity).
6. The 26S Proteasome Degradative Chamber
The 26S proteasome is a massive, cylindrical 2.5 MDa multicatalytic protease complex located in both the cytoplasm and nucleus of eukaryotic cells. It consists of one central 20S catalytic core particle capped at one or both ends by a 19S regulatory particle.
26S Proteasome Structure
+===========================+
| 19S Regulatory Cap | -> Recognises polyubiquitin, Cleaves Ub,
+---------------------------+ Unfolds protein using AAA+ ATPases
| 20S Core Particle |
| - α ring (7 subunits) | -> Structural gating, prevents entry
| - β ring (7 subunits) | -> Catalytic core (Thr active sites)
| - β ring (7 subunits) | -> Catalytic core (Thr active sites)
| - α ring (7 subunits) | -> Structural gating, prevents entry
+---------------------------+
| 19S Regulatory Cap |
+===========================+
6.1 The 20S Core Particle
The 20S core particle is a hollow cylinder formed by the stacking of four heptameric rings, exhibiting $\alpha_7\beta_7\beta_7\alpha_7$ stoichiometry:
- The Outer $\alpha$-Rings: Formed by seven structurally related but distinct $\alpha$-subunits ($\alpha_1$ to $\alpha_7$). These subunits are catalytically inactive; they act as a gated entryway, blocking unstructured proteins from entering the proteolytic chamber.
- The Inner $\beta$-Rings: Formed by seven distinct $\beta$-subunits ($\beta_1$ to $\beta_7$). These subunits line the interior of the cylinder, hosting active protease sites. The proteasome is an N-terminal nucleophile (Ntn) hydrolase that uses the hydroxyl group of an N-terminal threonine residue as the catalytic nucleophile.
Catalytic Specificities
Three specific $\beta$-subunits perform distinct proteolytic cleavages:
- $\beta_1$ Subunit (Caspase-like / Peptidyl-glutamyl peptide-hydrolysing activity): Cleaves peptide bonds on the carboxyl side of acidic residues (Asp, Glu).
- $\beta_2$ Subunit (Trypsin-like activity): Cleaves peptide bonds on the carboxyl side of basic residues (Arg, Lys).
- $\beta_5$ Subunit (Chymotrypsin-like activity): Cleaves peptide bonds on the carboxyl side of hydrophobic aromatic residues (Phe, Tyr, Trp).
6.2 The 19S Regulatory Particle
The 19S cap controls entry into the 20S core. It contains 19 individual protein subunits, organized into two sub-assemblies:
- The Base (9 Subunits): Contains six distinct AAA+ family ATPases (Rpt1–Rpt6) arranged in a hexameric ring. These ATPases use energy from ATP hydrolysis to physically unfold the target protein and thread it through the opened $\alpha$-ring into the 20S catalytic chamber.
- The Lid (10 Subunits): Contains receptors that recognize K48-linked polyubiquitin chains, and a specialized isopeptidase (deubiquitinating enzyme, DUB) that cleaves the polyubiquitin chain from the substrate, allowing the ubiquitin molecules to be recycled.
7. The N-End Rule Pathway
The in vivo half-life of intracellular proteins ranges from a few seconds to several days. This selective stability is governed by the N-end rule pathway, a specialized branch of the ubiquitin-proteasome system where the half-life of a protein is determined by the identity of its amino-terminal amino acid residue. This N-terminal signal is termed an N-degron.
UNSTABLE N-TERMINUS (e.g., Arg, Lys, Phe) ===> Recognised by N-recognins (E3) ===> Rapid UPS Degradation
STABLE N-TERMINUS (e.g., Ala, Gly, Val) ===> Ignored by N-recognins ===> Prolonged intracellular life
7.1 Structural Elements of the N-Degron
An N-degron consists of two essential features:
- A destabilising N-terminal amino acid residue.
- A nearby, accessible internal lysine residue within the unstructured segment of the target protein, which serves as the site for covalent polyubiquitination.
7.2 Destabilising Residues in Saccharomyces cerevisiae
In the yeast S. cerevisiae, the relationship between the N-terminal residue and protein stability is categorized into stabilizing and destabilizing classes:
| N-terminal Residue (X-$\beta$-gal) | In Vivo Half-Life in S. cerevisiae | Classification |
|---|---|---|
| Met, Gly, Ala, Ser, Thr, Val, Cys, Pro | $> 20$ hours (Up to $> 30$ hours) | Highly Stabilising |
| Tyr, His | 10 minutes (Tyr), $> 5$ hours (His) | Moderately Destabilising |
| Ile, Asp, Glu | 30 minutes | Highly Destabilising |
| Lys, Arg | 3 minutes (Lys), 2 minutes (Arg) | Extremely Destabilising (Type 1) |
| Phe, Leu, Trp, Tyr | 3 minutes | Extremely Destabilising (Type 2) |
| Asn, Gln | 3 minutes (Asn), 10 minutes (Gln) | Destabilising (Requires modification) |
7.3 Hierarchical Processing of N-Degrons
Eukaryotes categorize destabilizing N-terminal residues into three hierarchical tiers based on whether they are directly recognized by E3 ligases (N-recognins) or require enzymatic modification first:
Hierarchical Processing of N-Degrons
[ Tertiary Destabilising ] [ Secondary Destabilising ] [ Primary Destabilising ]
- Asparagine (Asn, N) - Aspartate (Asp, D) - Type 1 (Basic):
- Glutamine (Gln, Q) - Glutamate (Glu, E) Lys, Arg, His
- Oxidised Cysteine (Cys) - Type 2 (Hydrophobic):
| | Phe, Leu, Trp, Tyr, Ile
| N-terminal |
| Deamidase | Arginyl-tRNA-protein
v | transferase (ATE1)
[ Secondary Destabilising ] -----------------+
| |
+================================================================+
|
v
Recognised by N-recognin E3s
(e.g., UBR1 in yeast)
|
v
Polyubiquitinated & degraded
- Primary Destabilising Residues: Directly recognized by E3 ubiquitin ligases (N-recognins, e.g., UBR1 in yeast).
- Type 1 (Basic/Charged): Lysine (Lys), Arginine (Arg), and Histidine (His).
- Type 2 (Bulky Hydrophobic): Phenylalanine (Phe), Leucine (Leu), Tryptophan (Trp), Tyrosine (Tyr), and Isoleucine (Ile).
- Secondary Destabilising Residues: Aspartate (Asp) and Glutamate (Glu). To be recognized, they must first be modified by arginyl-tRNA-protein transferase (ATE1), which transfers a primary destabilizing arginine residue from Arg-tRNA to the free N-terminus of the protein.
- Tertiary Destabilising Residues: Asparagine (Asn) and Glutamine (Gln). They must first undergo deamidation of their side-chain amide groups by N-terminal deamidases to yield the secondary destabilizing residues Aspartate and Glutamate, which are then arginylated and degraded.
8. Protein Sequencing: N-Terminal Analysis and Edman Degradation
Determining the primary amino acid sequence of a purified protein is a core biochemical technique, accomplished by selective N-terminal labelling or sequential Edman degradation.
8.1 Reagents for N-Terminal Identification
Before sequencing a polypeptide, identifying the very first N-terminal residue can confirm protein purity and identity. This is achieved using two main electrophilic reagents:
Sanger’s Reagent (1-Fluoro-2,4-dinitrobenzene, FDNB)
Under mildly alkaline conditions (pH 9.5), the unprotonated N-terminal $\alpha$-amino group attacks the aromatic carbon of FDNB via nucleophilic aromatic substitution, releasing hydrofluoric acid (HF) and forming a yellow dinitrophenyl (DNP)-peptide derivative.
NO2 NO2
/ \ / O2N F + H2N--CH(R1)--CO--Peptide ===> O2N NH--CH(R1)--CO--Peptide + HF
\ / \ /
[ FDNB ] [ DNP-Peptide ]
|
+---> 6 M HCl Hydrolysis
(Cleaves all peptide bonds)
v
[ DNP-Amino Acid ] (Yellow)
+ Free Amino Acids
When the DNP-peptide is subjected to strong acid hydrolysis (6 M HCl), all internal peptide bonds are cleaved. However, the covalent carbon-nitrogen bond linking the dinitrophenyl group to the N-terminal amino acid remains intact. The yellow DNP-amino acid is then isolated and identified by chromatography.
Dansyl Chloride (5-Dimethylaminonaphthalene-1-sulfonyl chloride)
Operating via a similar mechanism, dansyl chloride reacts with the N-terminal amino group to form a dansyl-peptide intermediate. Following acid hydrolysis, the resulting dansyl-amino acid exhibits intense yellow-green fluorescence under ultraviolet light. This enables highly sensitive detection at nanomolar concentrations, far exceeding the sensitivity of Sanger’s reagent.
8.2 The Edman Degradation Sequencing Chemistry
Developed by Pehr Edman, this chemical method sequentially removes one amino acid residue at a time from the N-terminus of a peptide without cleaving the remaining peptide bonds. It proceeds through a three-stage reaction cycle:
[ STAGE 1: Coupling at pH 9 ]
PITC (Edman Reagent) + H2N--CH(R1)--CONH--Peptide
|
v
Phenylthiocarbamyl-Peptide (PTC-Peptide)
[ STAGE 2: Cyclisation / Cleavage using Anhydrous TFA ]
PTC-Peptide === TFA ===> Anilinothiazolinone-Amino Acid (ATZ-AA) + Shortened Peptide (R2)
(Extracted into organic solvent)
[ STAGE 3: Conversion using Aqueous Acid ]
ATZ-AA === H3O+ ===> Phenylthiohydantoin-Amino Acid (PTH-Amino Acid)
- Stable derivative, identified via chromatography
- Stage 1: Coupling (Alkaline Phase): The uncharged N-terminal amino group of the peptide reacts with phenylisothiocyanate (PITC, Edman’s reagent) in an alkaline solution (pH 9.0) to form a stable phenylthiocarbamyl-peptide (PTC-peptide).
- Stage 2: Cyclisation and Cleavage (Anhydrous Acid Phase): The PTC-peptide is treated with a strong anhydrous acid, such as trifluoroacetic acid (TFA). The sulfur atom of the phenylthiocarbamyl group attacks the carbonyl carbon of the first peptide bond, forming a cyclic intermediate. This selectively cleaves the first peptide bond, releasing the N-terminal residue as an anilinothiazolinone-amino acid (ATZ-amino acid) derivative. Importantly, because anhydrous acid is used, the remaining peptide bonds in the shortened polypeptide chain remain completely intact.
- Stage 3: Conversion (Aqueous Acid Phase): The unstable ATZ-amino acid is selectively extracted into an organic solvent and treated with dilute aqueous acid ($\text{H}_3\text{O}^+$). This prompts rearrangement into a highly stable phenylthiohydantoin-amino acid (PTH-amino acid). The PTH-amino acid is then identified by high-performance liquid chromatography (HPLC) or thin-layer chromatography, while the remaining shortened polypeptide is subjected to another round of Edman degradation.
9. Chemical and Enzymatic Cleavage Strategies
Standard Edman degradation can reliably sequence polypeptides of up to only 50 residues before signal loss occurs due to incomplete reactions. To sequence larger proteins, the polypeptide must first be cleaved into smaller peptides. This is achieved using sequence-specific endopeptidases or chemical reagents.
Peptide Cleavage Specificity Map
- - - NH--CH(Rn-1)--CO -- NH--CH(Rn)--CO - - -
| |
v v
(Amino Side) (Carboxyl Side)
9.1 Enzymatic Cleavage Specificity (Proteolytic Enzymes)
| Proteolytic Enzyme | Class | Cleavage Site Specificity | Key Restrictions |
|---|---|---|---|
| Trypsin | Endopeptidase | Carboxyl side of basic residues: Lys, Arg ($R_{n-1} = \text{Lys/Arg}$) | Will not cleave if next residue is Proline ($R_n = \text{Pro}$) |
| Chymotrypsin | Endopeptidase | Carboxyl side of aromatic/large hydrophobic residues: Tyr, Phe, Trp ($R_{n-1} = \text{Tyr/Phe/Trp}$) | Will not cleave if next residue is Proline ($R_n = \text{Pro}$) |
| Pepsin | Endopeptidase | Amino side of aromatic/hydrophobic residues: Tyr, Phe, Trp, Leu ($R_n = \text{Tyr/Phe/Trp/Leu}$) | Will not cleave if preceding residue is Proline ($R_{n-1} = \text{Pro}$) |
| Elastase | Endopeptidase | Carboxyl side of small neutral residues: Ala, Gly, Ser ($R_{n-1} = \text{Ala/Gly/Ser}$) | Will not cleave if next residue is Proline ($R_n = \text{Pro}$) |
| Thermolysin | Endopeptidase | Amino side of bulky hydrophobic residues: Leu, Ile, Val, Phe ($R_n = \text{Leu/Ile/Val/Phe}$) | Highly heat-stable metalloprotease |
| Carboxypeptidase A | Exopeptidase | Sequentially cleaves single residues from the C-terminus | Will not cleave C-terminal Pro, Lys, Arg |
| Carboxypeptidase B | Exopeptidase | Sequentially cleaves single residues from the C-terminus | Only cleaves C-terminal Arg, Lys |
| Carboxypeptidase C | Exopeptidase | Sequentially cleaves single residues from the C-terminus | Cleaves any C-terminal residue |
| Aminopeptidase M | Exopeptidase | Sequentially cleaves single residues from the free N-terminus | Cleaves all free N-terminal residues |
9.2 Chemical Cleavage Specificity
- Cyanogen Bromide (CNBr): Specifically cleaves peptide bonds on the carboxyl side of methionine residues. The reaction converts the C-terminal methionine residue of the released peptide into a cyclic peptidyl peptidyl-homoserine lactone.
- Hydroxylamine ($\text{NH}_2\text{OH}$): Specifically cleaves peptide bonds linking asparagine and glycine (Asn-Gly) residues.
10. Biophysical Methods for Disulfide Bond Analysis
Many native extracellular proteins contain disulfide bonds that link cysteine residues. For sequencing, these bonds must be identified or cleaved to allow the polypeptide chain to unfold completely.
DISULFIDE REDUCTION & ALKYLATION:
Protein-Cys-S-S-Cys-Protein
|
+---> Add DTT or β-mercaptoethanol (Reduces S-S to free -SH)
v
Protein-Cys-SH + HS-Cys-Protein
|
+---> Add Iodoacetate (Irreversible alkylation)
v
Protein-Cys-S-CH2-COO- + -OOC-CH2-S-Cys-Protein (Carboxymethyl-cysteines)
-----------------------------------------------------------------------------
PERFORMIC ACID OXIDATION:
Protein-Cys-S-S-Cys-Protein === Performic Acid ===> 2 x Protein-Cys-SO3-
(Stable Cysteic Acid residues,
completely blocks S-S reform)
- Performic Acid Oxidation: Treatment of a protein with performic acid oxidizes all disulfide bonds and free sulfhydryl groups into highly stable cysteic acid residues ($-\text{CH}_2-\text{SO}_3^-$). Because these cysteic acid side chains carry a strong negative charge, electrostatic repulsion prevents the disulfide bonds from reforming.
- Reduction and Alkylation: Alternatively, disulfide bonds are reduced to free sulfhydryl groups using excess dithiothreitol (DTT) or $\beta$-mercaptoethanol. Because these reduced sulfhydryl groups are highly reactive and will spontaneously oxidize back into disulfides in the presence of oxygen, they must be irreversibly blocked. This is achieved by adding iodoacetate, which alkylates the free sulfhydryl groups to form stable carboxymethyl-cysteine residues.
11. Analytical Protein Assays and Quantification Methods
Determining protein concentration in an unknown sample is a fundamental biochemical task, utilizing either direct UV spectroscopy or dye-binding colorimetric assays.
11.1 Ultraviolet Spectroscopy (Direct Quantification)
Proteins absorb ultraviolet light in the range of 190 nm to 300 nm through two distinct biophysical mechanisms:
- Peptide Backbone Absorption (190 nm – 210 nm): The peptide amide bond absorbs strongly in the far-UV spectrum ($205 \text{ nm}$). This assay is extremely sensitive, but is prone to interference from many buffer salts and solvents.
- Aromatic Side Chain Absorption (280 nm): The aromatic amino acids tryptophan (Trp, W) and tyrosine (Tyr, Y) exhibit strong UV absorption in the near-UV spectrum, peaking at $280 \text{ nm}$. Tryptophan’s molar absorptivity at $280 \text{ nm}$ is approximately four times greater than that of tyrosine. Phenylalanine absorbs weakly at $257 \text{ nm}$ and does not contribute significantly at $280 \text{ nm}$.
- The Nucleic Acid Correction: Nucleic acids absorb strongly at $260 \text{ nm}$ (due to purine and pyrimidine rings), which can skew protein measurements. To calculate protein concentration in samples contaminated with nucleic acids, the absorbance at both 280 nm and 260 nm is measured and corrected using the Warburg-Christian equation:
$$ \text{Protein Concentration (mg/mL)} = 1.55 \times A_{280} – 0.76 \times A_{260} $$
11.2 Classical Colorimetric Assays
When direct UV measurements are impractical due to buffer interference, colorimetric assays are used:
- The Biuret Method: In a strongly alkaline solution, divalent copper ions ($\text{Cu}^{2+}$) coordinate with the nitrogen atoms of four adjacent peptide bonds, forming a purple-coloured coordination complex. This assay is highly specific for proteins, but requires relatively high protein concentrations (low sensitivity).
- The Lowry Method (Folin Assay): This method combines the Biuret copper-coordination reaction with the reduction of the Folin-Ciocalteu reagent (phosphomolybdate-phosphotungstate). The copper-treated peptide backbone, along with tyrosine and tryptophan residues, reduces the Folin reagent to a deep blue heteropolymolybdenum complex. This assay is highly sensitive but sensitive to interference from reducing agents, detergents, and common laboratory buffers.
- The Bradford Assay (Coomassie Dye-Binding): This rapid, highly sensitive assay relies on the binding of Coomassie Brilliant Blue G-250 dye to proteins, particularly basic (Lys, Arg, His) and aromatic residues. Upon binding, the dye transitions from its protonated red/brown state (absorbing at $465 \text{ nm}$) to its stable, unprotonated blue anionic state, shifting its absorption maximum to $595 \text{ nm}$. The change in absorbance at $595 \text{ nm}$ is directly proportional to protein concentration.
Coomassie G-250 (Unbound, Acidic pH) =========> Coomassie G-250 (Bound to Protein) Red-brown anionic form Anionic blue form Absorbance max: 465 nm Absorbance max: 595 nm - The Ninhydrin Test (Amino Acid Identification): Ninhydrin (triketohydrindene hydrate) is a powerful oxidizing agent that reacts with the free $\alpha$-amino group of amino acids or peptides. The reaction proceeds through oxidative deamination, releasing ammonia ($\text{NH}_3$), carbon dioxide ($\text{CO}_2$), and an aldehyde. The released ammonia then condenses with one molecule of reduced ninhydrin (hydrindantin) and one molecule of oxidized ninhydrin to form a deep blue-purple complex known as Ruhemann’s purple, which absorbs strongly at $570 \text{ nm}$.
- The Proline Exception: Because proline is a secondary cyclic amine (imino acid), it cannot undergo oxidative deamination. Instead, it condenses with ninhydrin to form a distinct yellow-coloured product, absorbing at $440 \text{ nm}$.
12. Quantitative Protein Purification and Specific Activity
Isolating a single protein of interest from a complex cell lysate requires a series of purification steps based on size, charge, solubility, or binding affinity. To monitor the success of a purification protocol, both total protein concentration and the enzymatic or biological activity of the target protein must be quantified at each step.
12.1 Mathematical Formulas for Purification Analysis
- Total Protein (mg): The total mass of all protein in the fraction:
$$ \text{Total Protein (mg)} = \text{Protein Concentration (mg/mL)} \times \text{Total Volume (mL)} $$ - Total Activity (Units): The total catalytic capability of the target enzyme:
$$ \text{Total Activity (U)} = \text{Enzyme Activity Concentration (U/mL)} \times \text{Total Volume (mL)} $$ - Specific Activity (U/mg): A measure of enzyme purity, representing the units of enzyme activity per milligram of total protein:
$$ \text{Specific Activity} = \frac{\text{Total Activity (Units)}}{\text{Total Protein (mg)}} $$
As a protein is purified and contaminating proteins are removed, the Specific Activity must increase, reaching a maximum when the target protein is 100% pure. - Yield (%): The percentage of the starting enzyme activity recovered after a step:
$$ \text{Yield (\%)} = \left( \frac{\text{Total Activity in Current Step (U)}}{\text{Total Activity in Initial Crude Extract (U)}} \right) \times 100\% $$ - Purification Level (Fold): The increase in specific activity (purity) relative to the initial crude extract:
$$ \text{Purification Level (Fold)} = \frac{\text{Specific Activity in Current Step (U/mg)}}{\text{Specific Activity in Initial Crude Extract (U/mg)}} $$
13. Fully-Worked Biochemical Analysis Problems
Below are exhaustive solutions to the sample problems presented in the textbook material.
Problem 1: Cleavage Site Selectivity of Trypsin
Question: Which peptide bond(s) marked as a, b, c, d, and e will be broken when the following oligopeptide is treated with trypsin at pH 7.0?
$$ \text{Lys} \overset{\text{a}}{\sim} \text{Arg} \overset{\text{b}}{\sim} \text{Pro} \overset{\text{c}}{\sim} \text{Lys} \overset{\text{d}}{\sim} \text{Arg} \overset{\text{e}}{\sim} \text{Gly} $$
Step-by-Step Resolution:
- Recall Trypsin’s Specificity: Trypsin is an endopeptidase that specifically cleaves peptide bonds on the carboxyl side of basic amino acids, namely Lysine (Lys) and Arginine (Arg).
- Identify Potential Cleavage Sites:
- Bond a: Located on the carboxyl side of Lys1. This is a potential cleavage site.
- Bond b: Located on the carboxyl side of Arg2. This is a potential cleavage site.
- Bond c: Located on the carboxyl side of Pro3. Proline is not basic; this bond cannot be cleaved by trypsin.
- Bond d: Located on the carboxyl side of Lys4. This is a potential cleavage site.
- Bond e: Located on the carboxyl side of Arg5. This is a potential cleavage site.
- Apply Restrictions: Trypsin cannot cleave a Lys or Arg peptide bond if the subsequent amino acid residue ($R_n$) is Proline.
- For Bond a, the subsequent residue is Arg2. Cleavage is permitted.
- For Bond b, the subsequent residue is Pro3. Cleavage is completely blocked by the proline residue.
- For Bond d, the subsequent residue is Arg5. Cleavage is permitted.
- For Bond e, the subsequent residue is Gly6. Cleavage is permitted.
- Final Deduction: Trypsin will cleave peptide bonds a, d, and e. (Note: If the starting peptide is cleaved sequentially, initial cleavage at a splits Lys1 from the peptide, leaving Arg2 at the new N-terminus, but because Pro3 is adjacent, the bond b remains completely intact).
Problem 2: Sequencing by Peptide Fragment Overlap
Question: A polypeptide was cleaved into two smaller peptides by Cyanogen Bromide (CNBr), and into two different peptides by trypsin. The isolated fragments have the following sequences:
- CNBr 1: $ \text{Gly-Thr-Lys-Ala-Glu} $
- CNBr 2: $ \text{Ser-Met} $
- Trypsin 1: $ \text{Ser-Met-Gly-Thr-Lys} $
- Trypsin 2: $ \text{Ala-Glu} $
Determine the complete sequence of the parent peptide.
Step-by-Step Resolution:
- Analyze Cleavage Specificities:
- CNBr cleaves at the carboxyl side of Methionine (Met). Therefore, any peptide ending in Met (or homoserine lactone) must be an internal cleavage product, while the peptide lacking Met must represent the C-terminus. Thus, CNBr 2 ($\text{Ser-Met}$) is positioned before CNBr 1 ($\text{Gly-Thr-Lys-Ala-Glu}$).
- Trypsin cleaves on the carboxyl side of Lys and Arg. Therefore, Trypsin 1 ($\text{Ser-Met-Gly-Thr-Lys}$) must precede Trypsin 2 ($\text{Ala-Glu}$).
- Align the Overlapping Fragments: We arrange the fragments in an overlapping set to align matching sequences:
CNBr 2: Ser-Met Trypsin 1: Ser-Met-Gly-Thr-Lys CNBr 1: Gly-Thr-Lys-Ala-Glu Trypsin 2: Ala-Glu - Deduce Parent Sequence: By reading the aligned consensus sequence from N-terminus to C-terminus, we obtain:
$$ \text{Ser-Met-Gly-Thr-Lys-Ala-Glu} $$
Problem 3: Enkephalin Peptide Net Charge and Digestion Analysis
Question: Enkephalins are naturally occurring pentapeptides that act as endogenous opiates. Suppose an enkephalin has the following primary sequence:
$$ \text{Tyr-Gly-Gly-Phe-Ala-Met} $$
(Note: While classical enkephalins are pentapeptides, this synthetic analogue is a hexapeptide containing an additional C-terminal Methionine).
- Part a: What is the net charge of this peptide at pH 1.0 and pH 7.0?
- Part b: Indicate the number of bonds that could be hydrolysed (if any) by digestion with:
- Trypsin
- Chymotrypsin
- Cyanogen Bromide (CNBr)
Step-by-Step Resolution for Part a:
- Identify Ionisable Groups in $\text{Tyr-Gly-Gly-Phe-Ala-Met}$:
- The N-terminal $\alpha$-amino group of Tyrosine ($\text{p}K_{\text{a}} \approx 9.0$).
- The C-terminal $\alpha$-carboxyl group of Methionine ($\text{p}K_{\text{a}} \approx 2.0$).
- The phenolic hydroxyl group on the side chain of Tyrosine ($\text{p}K_{\text{R}} \approx 10.1$).
- All other residues (Gly, Gly, Phe, Ala, Met) have completely non-ionisable side chains.
- Calculate Net Charge at pH 1.0:
At pH 1.0 (strongly acidic, $\text{pH} < \text{p}K_{\text{a}}$ of all groups):- N-terminal $\alpha$-amino group is fully protonated: $-\text{NH}_3^+$ (Charge = $+1$).
- C-terminal $\alpha$-carboxyl group is fully protonated: $-\text{COOH}$ (Charge = $0$).
- Tyrosine side chain is fully protonated: $-\text{OH}$ (Charge = $0$).
- Net Charge at pH 1.0 = $+1$
- Calculate Net Charge at pH 7.0:
At pH 7.0 (neutral, $\text{p}K_{\alpha\text{-carboxyl}} < \text{pH} < \text{p}K_{\alpha\text{-amino}}, \text{p}K_{\text{R}}$):- N-terminal $\alpha$-amino group remains protonated: $-\text{NH}_3^+$ (Charge = $+1$).
- C-terminal $\alpha$-carboxyl group is deprotonated: $-\text{COO}^-$ (Charge = $-1$).
- Tyrosine side chain remains protonated: $-\text{OH}$ (Charge = $0$).
- Net Charge at pH 7.0 = $0$ (Zwitterionic state)
Step-by-Step Resolution for Part b:
- Digestion with Trypsin: Trypsin cleaves only on the carboxyl side of Lys and Arg. Since this hexapeptide contains neither Lysine nor Arginine, Trypsin will cleave 0 bonds.
- Digestion with Chymotrypsin: Chymotrypsin cleaves on the carboxyl side of aromatic residues: Tyrosine (Tyr) and Phenylalanine (Phe).
- Site 1: Carboxyl side of Tyr1. This peptide bond ($\text{Tyr-Gly}$) will be cleaved.
- Site 2: Carboxyl side of Phe4. This peptide bond ($\text{Phe-Ala}$) will be cleaved.
- Therefore, Chymotrypsin will cleave exactly 2 bonds, yielding three fragments: $\text{Tyr}$, $\text{Gly-Gly-Phe}$, and $\text{Ala-Met}$.
- Digestion with Cyanogen Bromide (CNBr): CNBr cleaves specifically on the carboxyl side of Methionine (Met) residues.
- In this sequence, Met6 is the absolute C-terminal amino acid.
- Because Met6 is at the C-terminus, there is no peptide bond adjacent to its carboxyl group to cleave.
- Therefore, CNBr will cleave 0 bonds.
Problem 4: Specific Activity and Purification Yield Analysis
Question: Calculate the Specific Activity and Purification Level (Fold) at each step of the purification protocol detailed in the table below:
| Step | Purification Procedure | Total Protein (mg) | Total Activity (Units) |
|---|---|---|---|
| 1 | Crude Cell Extract | 20,000 | 4,000,000 |
| 2 | Salt Precipitation | 5,000 | 3,000,000 |
| 3 | DEAE-Cellulose Chromatography | 1,500 | 1,000,000 |
| 4 | Size-Exclusion Chromatography | 500 | 750,000 |
| 5 | Affinity Chromatography | 45 | 675,000 |
Step-by-Step Calculations:
Step 1: Crude Cell Extract
$$ \text{Specific Activity} = \frac{4,000,000 \text{ Units}}{20,000 \text{ mg}} = 200 \text{ U/mg} $$
$$ \text{Purification Level} = \frac{200 \text{ U/mg}}{200 \text{ U/mg}} = 1.0\text{-fold} \text{ (Starting Reference)} $$
$$ \text{Yield} = \left( \frac{4,000,000 \text{ U}}{4,000,000 \text{ U}} \right) \times 100\% = 100\% $$
Step 2: Salt Precipitation
$$ \text{Specific Activity} = \frac{3,000,000 \text{ Units}}{5,000 \text{ mg}} = 600 \text{ U/mg} $$
$$ \text{Purification Level} = \frac{600 \text{ U/mg}}{200 \text{ U/mg}} = 3.0\text{-fold} $$
$$ \text{Yield} = \left( \frac{3,000,000 \text{ U}}{4,000,000 \text{ U}} \right) \times 100\% = 75\% $$
Step 3: DEAE-Cellulose Chromatography (Ion Exchange)
$$ \text{Specific Activity} = \frac{1,000,000 \text{ Units}}{1,500 \text{ mg}} \approx 666.67 \text{ U/mg} $$
$$ \text{Purification Level} = \frac{666.67 \text{ U/mg}}{200 \text{ U/mg}} \approx 3.33\text{-fold} $$
$$ \text{Yield} = \left( \frac{1,000,000 \text{ U}}{4,000,000 \text{ U}} \right) \times 100\% = 25\% $$
Step 4: Size-Exclusion Chromatography (Gel Filtration)
$$ \text{Specific Activity} = \frac{750,000 \text{ Units}}{500 \text{ mg}} = 1,500 \text{ U/mg} $$
$$ \text{Purification Level} = \frac{1,500 \text{ U/mg}}{200 \text{ U/mg}} = 7.5\text{-fold} $$
$$ \text{Yield} = \left( \frac{750,000 \text{ U}}{4,000,000 \text{ U}} \right) \times 100\% = 18.75\% $$
Step 5: Affinity Chromatography
$$ \text{Specific Activity} = \frac{675,000 \text{ Units}}{45 \text{ mg}} = 15,000 \text{ U/mg} $$
$$ \text{Purification Level} = \frac{15,000 \text{ U/mg}}{200 \text{ U/mg}} = 75.0\text{-fold} $$
$$ \text{Yield} = \left( \frac{675,000 \text{ U}}{4,000,000 \text{ U}} \right) \times 100\% = 16.875\% $$
Compiled Purification Summary Table:
| Step | Procedure | Total Protein (mg) | Total Activity (Units) | Specific Activity (U/mg) | Yield (%) | Purification Level (Fold) |
|---|---|---|---|---|---|---|
| 1 | Crude Extract | 20,000 | 4,000,000 | 200 | 100% | 1.0 (Reference) |
| 2 | Salt PPT | 5,000 | 3,000,000 | 600 | 75% | 3.0 |
| 3 | DEAE-Cellulose | 1,500 | 1,000,000 | 667 | 25% | 3.3 |
| 4 | Size-Exclusion | 500 | 750,000 | 1,500 | 18.8% | 7.5 |
| 5 | Affinity Chrom. | 45 | 675,000 | 15,000 | 16.9% | 75.0 |
Conclusion: The affinity chromatography step provides the highest increase in purity, bringing the enzyme to 75-fold purity compared to the starting material, while retaining 16.9% of the starting catalytic activity.
