Chapter 1: Foundations of Transcription and the Transcription Unit
1.1 The Chemistry and Logic of Transcription
Transcription is the template-directed enzymatic process of forming an RNA transcript from a DNA template. This process is catalyzed and rigorously scrutinized by RNA polymerases, utilizing complementary base pairing rules. Synthesis occurs strictly unidirectionally in the 5′ → 3′ direction, where each incoming ribonucleoside triphosphate (NTP) is appended to the free 3′-OH group of the nascent RNA chain.
While chemically and enzymatically similar to DNA replication, transcription possesses several defining characteristics that distinguish it from replication:
- Substrate Requirement: Transcription utilizes ribonucleoside triphosphates (ATP, GTP, UTP, CTP) containing ribose sugars, rather than deoxyribonucleoside triphosphates (dATP, dGTP, dTTP, dCTP) containing deoxyribose sugars.
- Primer Independence: Unlike DNA polymerases, which require a pre-existing 3′-OH primer, RNA polymerases are capable of initiating RNA synthesis de novo on a DNA template without the need for a primer.
- Selectivity: Replication is an all-or-nothing process where the entire genome is copied exactly once per cell cycle. In contrast, transcription is highly selective; only specific segments of the genome (transcription units) are transcribed at any given time, depending on cellular requirements.
- Template Asymmetry: During DNA replication, both strands of the parental DNA double helix act as templates. During transcription, only one of the two DNA strands serves as the template for RNA synthesis within any given region.
- Unidirectionality: Transcription proceeds unidirectionally along the template, transcribing only a single strand of DNA in the 3′ → 5′ direction to yield a 5′ → 3′ RNA transcript.
1.2 The Transcription Unit and Strand Nomenclature
A transcription unit is defined as a segment of DNA transcribed into a single RNA molecule, beginning at a transcription start site (TSS) and ending at a transcription termination site (TTS).
- Monocistronic Transcription Units: Typical of eukaryotes, these carry the information for a single gene, translating into a single polypeptide. Eukaryotic monocistronic units can be simple (processed to yield a single type of mRNA and polypeptide) or complex (capable of alternative processing, such as alternative splicing, to yield multiple different mRNAs and polypeptides from a single primary transcript).
- Polycistronic Transcription Units: Typical of prokaryotes, these comprise a set of adjacent genes transcribed as a single continuous RNA unit, containing multiple translation initiation and termination signals to encode several distinct polypeptides.
Strand Nomenclature:
- Template Strand (Antisense / Minus (-) Strand): The DNA strand that is physically read and transcribed by RNA polymerase in the 3′ → 5′ direction. It is complementary and antiparallel to the synthesized RNA transcript.
- Coding Strand (Sense / Plus (+) Strand): The DNA strand that is complementary to the template strand. Its sequence is identical to that of the synthesized RNA transcript (written in the 5′ → 3′ direction), except that thymine (T) in DNA is replaced by uracil (U) in RNA.
Numerical Coordinates:
- The first nucleotide of the transcribed sequence is designated as the transcription start site (TSS) and assigned the coordinate +1.
- Sequences preceding the TSS are designated as upstream and allotted negative numerical values (e.g., -10, -35), decreasing in the upstream direction. There is no coordinate 0.
- Sequences following the TSS are designated as downstream and allotted positive numerical values (e.g., +10, +100), increasing in the downstream direction.
Chapter 2: The Eubacterial RNA Polymerase Holoenzyme
2.1 Discovery and Subunit Composition
Eubacterial RNA polymerase was discovered in 1960 by Samuel B. Weiss and Jerard Hurwitz. In eubacteria (such as Escherichia coli), a single type of RNA polymerase is responsible for the synthesis of all cellular RNAs, including messenger RNA (mRNA), ribosomal RNA (rRNA), and transfer RNA (tRNA).
The complete, active enzyme is called the holoenzyme (α2ββ′ωσ), which has a molecular mass of approximately 465 kDa. It can be biochemically separated into two components: the core enzyme (α2ββ′ω) and the sigma factor (σ).
- Core Enzyme (α2ββ′ω): Retains full catalytic capability for RNA chain elongation but is completely incapable of recognizing promoters or initiating transcription at specific sites. It binds DNA non-specifically with high affinity.
- Sigma Factor (σ): A transiently associated subunit that reduces the core enzyme's affinity for non-promoter DNA sequences while dramatically increasing its affinity for specific promoter sequences. Thus, the sigma factor is strictly required for promoter recognition and transcription initiation.
2.2 Subunit Breakdown and Genetic Loci
The subunits of the eubacterial RNA polymerase core and holoenzyme are highly conserved and detailed below:
| Subunit | Gene | Molecular Weight | Stoichiometry | Key Biochemical Functions |
|---|---|---|---|---|
| Alpha (α) | rpoA | 36.5 kDa (each) | 2 (α2) | Scaffold for core enzyme assembly; promoter recognition of UP elements via the C-terminal domain (α-CTD); interaction with regulatory transcription factors. |
| Beta (β) | rpoB | 151 kDa | 1 | Forms the catalytic center; binds incoming ribonucleotides; involved in RNA chain elongation; target site for rifampicin inhibition. |
| Beta-prime (β′) | rpoC | 155 kDa | 1 | Forms the catalytic center; provides a structural cleft for binding the template DNA; coordinates the catalytic metal ions. |
| Omega (ω) | rpoZ | 10 kDa | 1 | Facilitates proper folding and assembly of the β′ subunit into the core enzyme; stabilizes the assembled core complex. |
| Sigma (σ) | rpoD (for σ70) | 70 kDa | 1 | Mediates promoter recognition by binding specifically to the -10 and -35 consensus elements; facilitates DNA melting to form the open complex. |
2.3 Viral and RNA-Dependent RNA Polymerases
While cellular transcription is template-dependent on DNA, certain genetic systems utilize alternative replication and transcription pathways:
- Viral RNA-Dependent RNA Polymerases (RNA Replicases): Some viruses, such as bacteriophages f2 and R17, contain single-stranded RNA genomes. These viruses replicate their genomes within host cells using RNA replicases (RNA-dependent RNA polymerases). These enzymes use viral RNA as a template to synthesize a complementary RNA strand in the 5′ → 3′ direction. The chemical mechanism is identical to DNA-dependent transcription, but RNA replicases are highly template-specific for their own viral RNAs and will not transcribe or replicate host cellular RNAs.
- Monomeric DNA-Dependent RNA Polymerases: Bacteriophages such as T7 encode their own single-chain monomeric RNA polymerases (~100 kDa) that are highly active and recognize specific, unique phage promoters. This contrasts with the multimeric, complex host RNA polymerases.
Chapter 3: Eukaryotic and Organellar RNA Polymerases
3.1 Nuclear RNA Polymerase Classification
Eukaryotes exhibit a division of labor among nuclear RNA polymerases, utilizing five distinct multimeric enzymes (Pol I, Pol II, Pol III, Pol IV, and Pol V) to transcribe different classes of genes.
- RNA Polymerase I (Pol I): Consists of 14 subunits. It is located in the nucleolus and is responsible for transcribing the single precursor gene for the major ribosomal RNAs, which is processed into the 28S, 18S, and 5.8S rRNAs. It is completely resistant to α-amanitin.
- RNA Polymerase II (Pol II): Consists of 12 subunits (RPB1 to RPB12). It transcribes all protein-coding genes into messenger RNA (mRNA), as well as most small nuclear RNAs (snRNAs), microRNAs (miRNAs), and long interspersed nuclear elements (LINEs). Pol II is extremely sensitive to α-amanitin.
- RNA Polymerase III (Pol III): Consists of 17 subunits. It transcribes small, structured non-coding RNAs, including transfer RNA (tRNA), the ribosomal 5S rRNA, the spliceosomal U6 snRNA, H1 RNA (the RNA component of RNase P involved in tRNA processing), MRP RNA (involved in rRNA processing), 7SL scRNA (the signal recognition particle RNA component), 7SK scRNA, and SINE transcripts. It exhibits moderate, species-specific sensitivity to α-amanitin.
- RNA Polymerase IV and V (Pol IV & Pol V): Plant-specific enzymes that evolved from duplications of RNA Polymerase II subunits. They are required for the biogenesis of small interfering RNAs (siRNAs) and are central to the siRNA-directed DNA methylation (RdDM) pathway, wherein siRNAs direct de novo cytosine methylation of complementary genomic DNA sequences to transcriptionally silence transposons and genes.
3.2 α-Amanitin Sensitivity
α-Amanitin is a highly toxic, cyclic octapeptide produced by the poisonous mushroom Amanita phalloides (commonly known as the death cap or destroying angel). It binds specifically to a conserved pocket in the active site of eukaryotic RNA Polymerase II, physically blocking the translocation step during transcription elongation.
The differential sensitivity of eukaryotic RNA polymerases to α-amanitin is a critical biochemical tool for distinguishing their active products:
| Polymerase | Subunits | Primary Transcription Products | Sensitivity to α-Amanitin |
|---|---|---|---|
| RNA Pol I | 14 | Precursor rRNA (processed into 5.8S, 18S, 28S rRNAs) | Completely Resistant |
| RNA Pol II | 12 | mRNA, most snRNA, LINEs, miRNA | Highly Sensitive (inhibited by ~10-9 M or 1 μg/mL) |
| RNA Pol III | 17 | tRNA, 5S rRNA, U6 snRNA, H1 RNA, 7SL/7SK scRNA | Moderately Sensitive (inhibited by ~10-6 M or 10-100 μg/mL in vertebrates; resistant in yeast) |
3.3 Structural Hallmarks of RNA Polymerase II: The CTD
The largest subunit of eukaryotic RNA Polymerase II, RPB1, contains a unique structural feature at its carboxyl terminus known as the Carboxyl-Terminal Domain (CTD).
- The CTD consists of multiple tandem repeats of a consensus heptapeptide sequence:
Tyr1 - Ser2 - Pro3 - Thr4 - Ser5 - Pro6 - Ser7 - Vertebrate Pol II contains 52 of these heptad repeats, while budding yeast contains 26.
- The hydroxyl-bearing side chains of tyrosine (Tyr1), serine (Ser2, Ser5, Ser7), and threonine (Thr4) are targets for intensive reversible phosphorylation.
The phosphorylation state of the CTD acts as a "glycoprotein scaffold" or "CTD code" that coordinates transcription initiation, transition to elongation, and mRNA processing:
- Unphosphorylated CTD: Required for Pol II recruitment and assembly into the pre-initiation complex (PIC).
- Serine-5 (Ser5) Phosphorylation: Catalyzed by the kinase subunit of TFIIH. It marks the transition to transcription initiation and recruits the mRNA 5′ capping enzymes.
- Serine-2 (Ser2) Phosphorylation: Catalyzed by positive transcription elongation factor b (p-TEFb) during elongation. It marks highly processive elongation and recruits splicing machinery and 3′ polyadenylation/cleavage complexes.
3.4 Organellar RNA Polymerases
Eukaryotic cells also contain organellar genomes within mitochondria and chloroplasts, which are transcribed by specialized, organelle-resident RNA polymerases:
- Mitochondrial RNA Polymerase: A single-subunit, monomeric enzyme encoded by a nuclear gene. It has a molecular weight of approximately 100 kDa and structurally resembles the monomeric RNA polymerase of bacteriophage T7. It works in conjunction with mitochondrial transcription factors to transcribe the circular mitochondrial DNA.
- Chloroplast RNA Polymerases: Chloroplasts of higher plants utilize two distinct classes of RNA polymerases:
- Plastid-Encoded RNA Polymerase (PEP): A multimeric, eubacterial-type enzyme whose core subunits (α, β, β′, ω) are encoded by the chloroplast genome, while its σ-like promoter-recognition factors are encoded by nuclear genes. PEP primarily transcribes photosynthetic genes.
- Nuclear-Encoded RNA Polymerase (NEP): A monomeric, phage-like single-subunit enzyme encoded by nuclear genes. NEP primarily transcribes housekeeping genes within the chloroplast.
Chapter 4: Prokaryotic Transcription Initiation, Promoters, and Alternative Sigma Factors
4.1 Eubacterial Promoter Architecture
A promoter is a cis-acting, position-dependent DNA sequence required for the accurate and efficient initiation of transcription. It provides the binding site for the RNA polymerase holoenzyme. In eubacteria, promoters recognized by the primary sigma factor (σ70) share several conserved structural elements:
- The -10 Box (Pribnow Box): A 6-base-pair consensus sequence located approximately 10 nucleotides upstream of the TSS. Its consensus sequence is: 5′-TATAAT-3′. This region is highly AT-rich, which facilitates the initial melting of the double helix during the transition from a closed to an open complex.
- The -35 Box: A 6-base-pair consensus sequence located approximately 35 nucleotides upstream of the TSS. Its consensus sequence is: 5′-TTGACA-3′. This region serves as the primary initial anchoring site for the σ factor of the holoenzyme.
- The Spacer Region: A non-conserved sequence separating the -35 and -10 boxes. Its length is highly constrained, typically measuring 16 to 18 base pairs. Deviations from this optimal spacing significantly decrease promoter strength.
- The UP Element (Upstream Promoter Element): An AT-rich sequence found in highly active promoters (such as ribosomal RNA promoters) located between coordinates -40 and -60. It is recognized not by the sigma factor, but by the C-terminal domain of the RNA polymerase alpha subunit (α-CTD).
- The +1 Start Site: Typically a purine nucleotide (adenine or guanine).
4.2 Promoter Strength and Consensus Sequence Matching
Promoter "strength" is defined as the frequency of productive transcription initiations promoted per second. Strong promoters drive high levels of transcription and match the σ70 consensus sequences (5′-TTGACA-3′ and 5′-TATAAT-3′) very closely. Weak promoters contain nucleotide substitutions that deviate from the consensus sequence or have suboptimal spacer lengths, resulting in lower affinity for the holoenzyme and a lower rate of open complex formation.
4.3 Alternative Sigma Factors
E. coli contains multiple alternative sigma factors that are expressed or activated under specific environmental conditions. These alternative sigma factors redirect the single RNA polymerase core enzyme to distinct regulons by recognizing completely different promoter consensus sequences:
| Sigma Factor | Gene | Environmental/Physiological Role | -35 Consensus | -10 Consensus |
|---|---|---|---|---|
| σ70 (RpoD) | rpoD | Housekeeping / Exponential growth | 5′-TTGACA-3′ | 5′-TATAAT-3′ |
| σ54 (RpoN) | rpoN | Nitrogen assimilation / stress | 5′-TTGGCACA-3′ (at -24) | 5′-TTGCA-3′ (at -12) |
| σ38 (RpoS) | rpoS | Stationary phase / General stress | 5′-CCGGCG-3′ | 5′-TATACT-3′ |
| σ32 (RpoH) | rpoH | Heat shock response / chaperones | 5′-TNTCNCCTTGAA-3′ | 5′-CCCCATNT-3′ |
| σ28 (RpoF) | fliA | Flagellar synthesis and chemotaxis | 5′-TAAA-3′ | 5′-GCCGATAA-3′ |
4.4 Stepwise Initiation Pathway
Prokaryotic transcription initiation is a highly regulated, multi-step kinetic pathway:
- Forming the Closed Binary Complex: The RNA polymerase holoenzyme binds to the double-stranded DNA promoter. The σ factor coordinates the recognition of the -35 and -10 boxes. In this state, the DNA remains completely double-stranded. The holoenzyme covers a large footprint on the DNA, spanning from approximately -55 to +20.
- Forming the Open Binary Complex: The bound holoenzyme undergoes a conformational change to melt approximately 12 to 14 base pairs of DNA, spanning from the -10 box to coordinate +2. This localized unwinding creates the transcription bubble. Unwinding is initiated within the AT-rich -10 box (Pribnow box) due to the lower thermodynamic stability of A-T base pairs.
- Forming the Ternary Complex: The open complex initiates RNA synthesis de novo. The first initiating nucleotide is a purine nucleoside triphosphate (pppG or pppA) that base-pairs with the +1 position of the template strand. It retains its 5′ triphosphate group (5′-pppG or 5′-pppA). The second NTP enters the active site, and the first phosphodiester bond is formed. This state, containing DNA, template-bound RNA, and enzyme, is called the ternary complex.
- Abortive Initiation: Before entering a highly processive elongation mode, the RNA polymerase repeatedly synthesizes and releases short transcripts ranging from 2 to 9 nucleotides in length without leaving the promoter. This occurs because the exit channel for the nascent RNA is physically blocked by the σ factor (specifically region 3.2).
- Promoter Clearance: Once the polymerase succeeds in synthesizing a transcript longer than 10 nucleotides, the growing RNA chain physically displaces the blocking domain of the σ factor. This triggers a massive conformational change: the affinity of the enzyme for the promoter is drastically reduced, the σ factor is released from the core enzyme, and the core polymerase "clears" the promoter to transition into highly processive transcription elongation.
Chapter 5: Prokaryotic Elongation, Topological Stress, and Transcription Inhibitors
5.1 Biophysics of Elongation
During transcription elongation, the core RNA polymerase moves rapidly along the template DNA strand, unwinding the DNA helix ahead of its path and rewinding it behind.
- The Transcription Bubble: The single-stranded transcription bubble is maintained at a constant length of approximately 12 to 14 base pairs during elongation.
- RNA-DNA Hybrid: Within the transcription bubble, the growing 3′ end of the nascent RNA remains base-paired with the DNA template strand, forming an RNA-DNA hybrid helix of approximately 8 to 9 base pairs in length.
- Catalytic Mechanism: The catalytic center of the active site utilizes a two-metal-ion mechanism coordinating two divalent cations (Mg2+ or Zn2+) via highly conserved aspartate residues. One metal ion activates the 3′-OH group of the growing RNA chain for nucleophilic attack, while the second stabilizes the negative charges of the incoming NTP's triphosphate group and facilitates the exit of the pyrophosphate (PPi) leaving group.
- Elongation Rate: Under optimal conditions at 37°C, the E. coli RNA polymerase core transcribes at an average rate of 40 nucleotides per second.
5.2 Topological Stress and Supercoiling
As the RNA polymerase complex moves along the DNA, it cannot rotate around the helical axis of the DNA template. This generates severe topological stress in the DNA:
- Positive Supercoiling Ahead: The forward movement of the polymerase overwinds the DNA ahead of the transcription bubble, generating positive supercoils.
- Negative Supercoiling Behind: The movement of the polymerase leaves underwound DNA behind the transcription bubble, generating negative supercoils.
- Topological Resolution: To prevent transcriptional arrest due to excessive topological tension:
- DNA Gyrase (Type II Topoisomerase) actively cuts and reseals DNA ahead of the polymerase to relieve positive supercoils.
- Topoisomerase I (Type I Topoisomerase) acts behind the polymerase to relieve negative supercoils.
5.3 Pharmacological and Chemical Inhibitors
Transcription elongation is a major target for various naturally occurring toxins and clinical antibiotics:
- Rifampicin: A semisynthetic derivative of the rifamycin family. It binds highly specifically to a pocket within the eubacterial β subunit (rpoB) near the catalytic center. Rifampicin does not inhibit the binding of DNA or the formation of the first phosphodiester bond. Instead, it acts as a steric roadblock: when the growing RNA chain reaches 2 to 3 nucleotides in length, it physically collides with the bound rifampicin molecule, blocking further elongation and preventing promoter clearance.
- Actinomycin D: An intercalating antibiotic containing a planar phenoxazone ring and two cyclic peptide lactone chains. The phenoxazone ring intercalates non-covalently between adjacent GC base pairs in double-stranded DNA, while the peptide chains fit snugly within the minor groove. This severely distorts the DNA double helix and blocks the movement of both prokaryotic and eukaryotic RNA polymerases along the template.
- Cordycepin (3′-deoxyadenosine): An analog of adenosine that lacks the 3′-OH group. Inside the cell, cordycepin is triphosphorylated to cordycepin triphosphate and is readily incorporated into growing RNA chains by RNA polymerase. However, because it lacks a 3′-OH group, it is impossible to form a phosphodiester bond with the next incoming nucleotide, resulting in immediate and irreversible transcription chain termination.
Chapter 6: Prokaryotic Transcription Termination (Intrinsic vs. Rho-Dependent)
To release the finished transcript and recycle the RNA polymerase core, transcription must terminate at specific sequences called terminators. Eubacteria utilize two distinct mechanical pathways for transcription termination: intrinsic termination and Rho-dependent termination.
6.1 Intrinsic (Rho-Independent) Termination
Intrinsic termination requires no accessory protein factors. It relies entirely on the primary sequence and secondary structure of the synthesized RNA transcript itself. Intrinsic terminators consist of two conserved regions:
- A GC-Rich Inverted Repeat Sequence: When transcribed, these self-complementary sequences rapidly fold into a highly stable hairpin (stem-loop) structure in the nascent RNA.
- A Poly-U Tract: A sequence of 7 to 9 consecutive uracil residues located immediately downstream of the inverted repeat (transcribed from an AT-rich template run).
- Mechanism: The formation of the hairpin loop occurs immediately as the RNA exits the polymerase. This bulky hairpin physically stalls the RNA polymerase and disrupts the active site. At the same time, the RNA-DNA hybrid remaining in the active site consists solely of weak rU-dA base pairs (which have only two hydrogen bonds per pair and represent the weakest thermodynamic pairing of any nucleic acid hybrid). The combined mechanical stress of the hairpin formation pulling the RNA upstream (shearing model) or destabilizing the active site pocket (allosteric model) causes the weak rU-dA hybrid to dissociate, releasing the RNA transcript and dismantling the transcription bubble.
6.2 Rho-Dependent Termination
Rho-dependent termination requires an essential protein factor called Rho (ρ). Rho is a specialized, ring-shaped hexameric protein (~275 kDa) with ATP-dependent RNA helicase activity.
- The rut Site (Rho Utilization Site): Rho recognizes and binds to a specific single-stranded sequence on the emerging RNA transcript called the rut site. The rut site is typically 50 to 90 nucleotides long, rich in cytidine residues, and poor in guanosine residues.
- Mechanism: Once bound to the rut site, Rho undergoes a conformational change to close its ring around the single-stranded RNA. It then translocates downstream along the transcript in the 5′ → 3′ direction, utilizing the energy of ATP hydrolysis. When the actively transcribing RNA polymerase encounters a downstream pausing signal (a sequence that temporarily stalls elongation), Rho catches up to the paused polymerase. Using its ATP-dependent RNA-DNA helicase activity, Rho unwinds the RNA-DNA hybrid helix within the active site, releasing the transcript and displacing the RNA polymerase core from the DNA.
Chapter 7: Eukaryotic Promoter Architecture (Pol I, Pol II, and Pol III)
Eukaryotic promoters are highly specialized to coordinate recruitment of distinct multi-protein transcription machineries.
7.1 RNA Polymerase I Promoter
The Pol I promoter is dedicated exclusively to transcribing the large ribosomal RNA precursor gene. It consists of two essential, GC-rich elements:
- The Core Promoter: Spans the transcription start site, extending from coordinate -45 to +20. It is essential for initiating transcription.
- The Upstream Control Element (UCE): A sequence of approximately 100 base pairs located further upstream, from coordinate -180 to -100. It significantly enhances the efficiency of transcription initiation.
7.2 RNA Polymerase III Promoter
Pol III transcription is unique because many of its promoters are internal promoters located entirely downstream of the TSS within the transcribed region of the gene:
- Type I Internal Promoter (e.g., 5S rRNA gene): Contains two internal conserved sequence blocks: Box A (located between +50 and +70) and Box C (located between +80 and +90).
- Type II Internal Promoter (e.g., tRNA genes): Contains Box A (located between +10 and +30) and Box B (located between +50 and +70).
- Type III Upstream Promoter (e.g., U6 snRNA gene): Unlike the other types, this promoter is located upstream of the TSS and structurally resembles Pol II promoters. It contains:
- A TATA Box centered at coordinate -30.
- A Proximal Sequence Element (PSE) centered at coordinate -60.
- A Distal Sequence Element (DSE) centered between -244 and -214.
7.3 RNA Polymerase II Promoter
The Pol II promoter has a modular architecture, consisting of a core promoter (which recruits the general transcription factors) and various upstream regulatory elements (which modulate transcription frequency).
7.3.1 Core Promoter Elements
The core promoter spans approximately 40 to 60 base pairs around the TSS. It typically contains a combination of the following conserved sequence motifs:
- BRE (TFIIB Recognition Element): Located at -37 to -32. It is recognized and bound by TFIIB.
- TATA Box: Typically located at -31 to -26. Its consensus sequence is 5′-TATA(A/T)(A/T)(A/G)-3′. It is recognized and bound by the TATA-Binding Protein (TBP) subunit of TFIID.
- Inr (Initiator Element): Spans the transcription start site from -2 to +5. Its consensus sequence is Py-Py-A(+1)-N-(T/A)-Py-Py.
- MTE (Motif-Ten Element): Located downstream of the TSS at +18 to +27.
- DPE (Downstream Promoter Element): Located at +28 to +32. It is highly conserved in TATA-less promoters and is recognized by TBP-Associated Factors (TAFs) in TFIID.
- DCE (Downstream Core Element): Composed of three sub-elements that assist in recruiting TFIID.
7.3.2 Regulatory Promoter Elements
Located further upstream of the core promoter, these elements bind sequence-specific transcriptional activators:
- CAAT Box: Typically centered around coordinate -80. It binds the transcription factors NF-1 and NF-Y. It is orientation-independent.
- GC Box: Contains the consensus sequence 5′-GGGCGG-3′. It binds the activator protein Sp1 and is orientation-independent.
Chapter 8: Eukaryotic Pre-Initiation Complex Assembly and Transcription Factors
Eukaryotic RNA Polymerase II cannot recognize promoters or initiate transcription on its own. It requires the stepwise assembly of a group of accessory proteins called General Transcription Factors (GTFs) to form the Pre-Initiation Complex (PIC).
8.1 The General Transcription Factors (GTFs)
The properties and functions of the six essential GTFs required for Pol II transcription are detailed below:
| GTF | Subunit Composition | Major Biochemical and Structural Functions |
|---|---|---|
| TFIID | TBP (TATA-Binding Protein) + 14 TAFs (TBP-Associated Factors) | Initial Recognizer: TBP binds the TATA box in the minor groove, bending the DNA by 80° and widening the minor groove via β-sheet insertion. TAFs recognize and bind the Inr and DPE elements, and regulate TBP binding. |
| TFIIA | 3 subunits | Stabilizer: Binds to TFIID, stabilizing the TBP-DNA interaction and preventing the binding of inhibitory factors. |
| TFIIB | Single subunit | Bridge: Binds to TBP and the BRE promoter element. It recruits the Pol II-TFIIF complex and helps establish the correct spacing between the promoter and the active site. |
| TFIIF | 2 subunits | Recruiter / Escort: Binds directly to RNA Polymerase II, preventing it from binding to non-specific DNA sequences, and escorts it to the assembling PIC. |
| TFIIE | 2 subunits | Gatekeeper: Recruits TFIIH to the complex and regulates its catalytic activities (helicase and kinase). |
| TFIIH | 10 subunits | Engine & Kinase: Contains two ATP-dependent DNA helicases (XPB and XPD) that melt the promoter DNA to form the open complex, and a cyclin-dependent kinase (CDK7) that phosphorylates Ser5 of the Pol II CTD to trigger promoter clearance. |
8.2 Stepwise Assembly Pathway of the PIC
The assembly of the Pre-Initiation Complex follows a strict, sequential order of protein recruitment:
- TFIID Binding: TFIID binds to the TATA box via TBP. This deforms the DNA helix, creating a highly bent platform.
- TFIIA and TFIIB Recruitment: TFIIA binds to stabilize TFIID. TFIIB then binds adjacent to TBP, recognizing the BRE and extending downstream.
- RNA Pol II - TFIIF Recruitment: RNA Polymerase II, pre-assembled with TFIIF, is recruited to the promoter, bridging with TFIIB.
- TFIIE and TFIIH Recruitment: TFIIE binds, creating the docking site for TFIIH. This completes the assembly of the closed pre-initiation complex.
- Promoter Melting: TFIIH uses its ATP-dependent DNA helicase activity to unwind the DNA near the transcription start site, converting the closed PIC into the open complex.
- CTD Phosphorylation and Promoter Clearance: The kinase subunit of TFIIH phosphorylates the serine-5 (Ser5) residues on the carboxyl-terminal domain (CTD) of the RPB1 subunit of Pol II. This phosphorylation disrupts the contacts between Pol II and the GTFs, allowing the polymerase to clear the promoter and initiate transcription.
8.3 Chromatin-specific Accessory Factors: The FACT Complex
In eukaryotes, genomic DNA is packaged into chromatin, which poses a severe physical barrier to transcription. RNA Polymerase II requires specialized elongation factors to navigate through nucleosomes:
- The FACT Complex (Facilitates Chromatin Transcription): A highly conserved, heterodimeric histone chaperone composed of two subunits: Spt16 and Pob3 (in yeast) or SSRP1 (in metazoans).
- Mechanism: FACT physically interacts with the nucleosome. During elongation, it selectively destabilizes and removes a single H2A-H2B heterodimer from the nucleosome octamer immediately ahead of the transcribing RNA Polymerase II. This temporarily unwinds the DNA from the nucleosome, allowing the polymerase to transcribe through. Once the polymerase has passed, FACT immediately chaperone-reassembles the H2A-H2B dimer back onto the remaining H3-H4 tetramer, restoring nucleosomal structure behind the transcription machinery. This prevents chromatin disassembly and suppresses aberrant transcription initiation from cryptic promoters within the gene body.
Chapter 9: Eukaryotic Transcription Elongation, Chromatin Remodeling, and Termination Models
9.1 Coupling Elongation with mRNA Processing
During eukaryotic transcription elongation, the phosphorylation state of the RNA Polymerase II CTD transitions from a Ser5-phosphorylated state to a Ser2-phosphorylated state. This transition is essential for coordinating transcription with co-transcriptional mRNA modifications:
- Capping: Ser5 phosphorylation recruits the capping enzyme complex to add the 7-methylguanosine cap to the 5′ end of the nascent pre-mRNA as soon as it emerges from the exit channel (around 20-30 nt).
- Splicing: Subsequent Ser2 phosphorylation during elongation recruits splicing factors to remove introns from the pre-mRNA.
- 3′ End Processing: High levels of Ser2 phosphorylation recruit the cleavage and polyadenylation specificity factors (CPSF and CstF) to process the 3′ end of the transcript.
9.2 Transcription Termination Models for RNA Polymerase II
Eukaryotic RNA Polymerase II does not terminate at a simple, defined sequence. Instead, termination is tightly coupled to the cleavage and polyadenylation of the mRNA 3′ end, which is directed by the polyadenylation signal sequence (5′-AAUAAA-3′). Two primary models describe how Pol II transcription is terminated:
9.2.1 The Torpedo Model
- Mechanism: As RNA Polymerase II transcribes past the poly(A) signal sequence (5′-AAUAAA-3′), the emerging RNA is recognized and cleaved at the polyadenylation site by endonuclease complexes. This cleavage releases the mature, upstream mRNA for polyadenylation and export.
- However, the polymerase remains bound to the template DNA and continues transcribing a downstream RNA fragment. Because this downstream fragment was generated by cleavage, it lacks a protective 5′ cap, leaving a free 5′-monophosphate group.
- This 5′-monophosphate is recognized by a highly processive, 5′-to-3′ exonuclease called Xrn2 (in mammals) or Rat1 (in yeast). The exonuclease loads onto the 5′ end of the transcript and rapidly degrades it, moving faster than the transcribing RNA Polymerase II. When the exonuclease catches up to the polymerase, it physically collides with the active site, destabilizing the RNA-DNA hybrid and displacing the polymerase from the DNA template.
9.2.2 The Allosteric (Conformational) Model
- Mechanism: In this model, transcription of the poly(A) signal sequence (5′-AAUAAA-3′) triggers a major conformational change in the transcription elongation complex.
- The binding of cleavage and polyadenylation factors to the emerging poly(A) RNA sequence and to the CTD of Pol II leads to the dissociation of positive elongation factors (such as elongation-stimulating proteins) and the recruitment of termination factors.
- This conformational change significantly destabilizes the RNA Polymerase II complex, reducing its processivity and causing it to spontaneously dissociate from the DNA template shortly after transcribing the poly(A) signal.
Chapter 10: Activators, Co-activators, and Chromatin Regulation
Eukaryotic transcription is regulated by sequence-specific transcription factors (activators) and co-regulatory complexes that modulate chromatin structure.
10.1 Transcriptional Activators: Domain Architecture
Transcriptional activators are sequence-specific DNA-binding proteins that stimulate the rate of transcription initiation. They have a modular domain structure, typically containing two essential, independent domains:
- DNA-Binding Domain (DBD): Recognizes and binds to specific cis-acting promoter or regulatory sequences (e.g., enhancers, UAS).
- Activation Domain (AD): Interacts with the general transcription machinery (such as TFIID, TFIIB, or Mediator) or recruits chromatin-modifying complexes to stimulate transcription. Activation domains are classified based on their amino acid composition:
- Acidic Domains: Rich in aspartate and glutamate residues (e.g., in yeast Gal4 or herpesvirus VP16).
- Glutamine-Rich Domains: Found in activators such as Sp1.
- Proline-Rich Domains: Found in activators such as CTF/NF-1.
10.2 Co-activator Complexes
Co-activators are proteins or multi-protein complexes that stimulate transcription but do not bind DNA directly. Instead, they act as bridges, coordinating interactions between sequence-specific activators and the general transcription machinery. Co-activators fall into two main functional classes:
10.2.1 Chromatin-Modifying Complexes
These complexes carry out covalent post-translational modifications of histone tails within nucleosomes:
- Histone Acetyltransferases (HATs): Catalyze the transfer of acetyl groups from acetyl-CoA to specific lysine residues on histone H3 and H4 tails. The addition of negatively charged acetyl groups neutralizes the positive charge of the lysine residues, weakening the electrostatic interaction between the histone tails and the negatively charged DNA backbone. This relaxes the chromatin structure, making the promoter DNA accessible to transcription factors.
- Histone Methyltransferases (HMTs): Add methyl groups to specific lysine or arginine residues, creating binding sites for downstream regulatory proteins.
10.2.2 ATP-Dependent Chromatin Remodeling Complexes
These are large, multi-subunit complexes (such as the SWI/SNF complex) that utilize the energy of ATP hydrolysis to mobilize, slide, evict, or restructure nucleosomes. This remodeling exposes promoter consensus elements (such as the TATA box) that were previously wrapped around nucleosomes.
10.3 The Yeast GAL1 Case Study
The regulation of the yeast GAL1 gene (encoding galactokinase) is a classic eukaryotic regulatory model:
- Gal4 Activator: A sequence-specific transcriptional activator that binds to the Upstream Activating Sequence (UAS) located approximately 275 base pairs upstream of the GAL1 core promoter. The UAS contains four conserved 17-bp sequences, each of which binds a dimer of Gal4.
- Domain Function: Gal4 binds DNA via a zinc-finger DNA-binding domain and stimulates transcription using an acidic activation domain. When active, Gal4 recruits chromatin remodeling complexes and HATs to open the GAL1 promoter, and directly recruits TFIIB and the Mediator complex to assemble the transcription machinery.
Chapter 11: Long-Range Regulatory Elements, Insulators, and Locus Control Regions
Eukaryotic genomes utilize complex, long-range regulatory architectures to control gene expression across large chromosomal domains.
11.1 Enhancers and Silencers
- Enhancers: Cis-acting DNA sequences, typically 100 to 300 base pairs in length, that strongly stimulate transcription from linked promoters. They have three defining characteristics:
- They function over extremely long distances, often located 100 kilobases (kb) or more upstream or downstream of the target promoter.
- They are orientation-independent; they function equally well in either forward or reverse orientation.
- They are position-independent; they can be located upstream of the promoter, downstream of the gene, or even within introns.
- Enhancers contain clusters of binding sites for multiple transcriptional activators. They communicate with target promoters via DNA looping, bringing the enhancer-bound activators into direct physical contact with the promoter-assembled PIC, a process facilitated by the Mediator complex.
- Silencers: Structurally similar to enhancers but bind transcriptional repressors to silence gene expression. Like enhancers, they are orientation- and position-independent.
11.2 Insulators (Boundary Elements)
Because enhancers can act over long distances and are non-specific for promoters, cells must prevent inappropriate cross-activation of neighboring genes. This is accomplished by insulators, specialized DNA boundary elements. Insulators are classified into two functional categories:
- Enhancer-Blockers: When positioned between an enhancer and a promoter, an enhancer-blocking insulator prevents the enhancer from activating that promoter. It does not inhibit the enhancer's activity on other promoters that are not separated by the insulator. This blocking is mediated by proteins (such as CTCF in vertebrates) that bind the insulator and form physical loop barriers, isolating the enhancer and promoter into distinct topological domains.
- Barrier Elements: These prevent the spreading of transcriptionally inactive heterochromatin into adjacent domains of active, euchromatin genes, thereby maintaining active chromatin boundaries.
11.3 Locus Control Regions (LCRs)
- Definition: A Locus Control Region (LCR) is a specialized group of cis-acting DNA elements required for the high-level, tissue-specific, copy-number-dependent transcription of an entire gene cluster or locus.
- The β-Globin Case Study: The human β-globin gene cluster on chromosome 11 encodes embryonic (ε), fetal (Gγ, Aγ), and adult (δ, β) globin genes, which are expressed sequentially during development.
- LCR Architecture: Located far upstream of the embryonic ε-globin gene, the β-globin LCR contains five distinct DNase I Hypersensitive Sites (HS1 to HS5) spread over approximately 10 kb of DNA. These sites bind erythroid-specific transcription factors.
- Mechanism: The LCR organizes the entire locus into an active chromatin hub. It physically interacts with individual globin gene promoters sequentially during development via DNA looping, activating them at specific developmental stages. It also serves as a barrier preventing the flanking heterochromatin from silencing the β-globin locus.
Chapter 12: Structural Motifs of DNA-Binding Proteins
Transcription factors must recognize and bind specific double-stranded DNA sequences with high affinity. They do this through highly conserved structural domains called DNA-binding motifs. Most of these motifs utilize a specific segment of the protein, typically an α-helix, that fits into the major groove of the DNA double helix to make sequence-specific hydrogen bonds with the edges of the base pairs.
12.1 Helix-Turn-Helix (HTH) Motif
- Structure: The HTH motif was the first DNA-binding structure identified. It is typically a small domain of approximately 20 amino acids. It consists of two hydrophobic α-helices separated by a tight β-turn (which frequently contains a glycine residue to facilitate the sharp bend).
- Mechanism: The second α-helix is called the recognition helix. It lies within the major groove of the DNA to make sequence-specific contacts. The first α-helix lies across the major groove, stabilizing the interaction.
- Homeodomain: An extended variation of the HTH motif found in developmental regulatory proteins (encoded by homeotic genes). The homeodomain is a highly conserved 60-amino-acid domain encoded by a 180-bp DNA sequence called the homeobox. It contains three α-helices: helices 2 and 3 form an HTH motif, with helix 3 serving as the recognition helix, while helix 1 projects into the minor groove to make additional contacts.
- Examples: E. coli Lac repressor, CAP (Catabolite Activator Protein), and bacteriophage Cro.
12.2 Helix-Loop-Helix (HLH) Motif
- Structure: The HLH motif mediates both sequence-specific DNA binding and protein dimerization. It consists of a short α-helix (15 to 16 amino acids) connected by a flexible, non-helical loop of variable length (12 to 28 amino acids) to a second, longer α-helix.
- Mechanism: The HLH motif forms stable homo- or heterodimers. Dimerization is mediated by the hydrophobic faces of the helices interacting with one another. The basic region at the N-terminal end of the recognition helix binds to the major groove of the DNA.
- Examples: MyoD (muscle-specific transcription factor) and Myc.
12.3 Leucine Zipper Motif
- Structure: The Leucine Zipper (bZip) motif is designed for protein dimerization and sequence-specific DNA binding. It consists of an amphipathic α-helix containing a leucine residue at every seventh position (a heptad repeat) along a stretch of 30 to 40 amino acids.
- Mechanism: Due to the helical turn of 3.6 residues per turn, the leucine residues project from the same face of the α-helix every two turns. Two monomeric helices wind around each other in a parallel orientation to form a coiled-coil structure, held together by hydrophobic interactions between the leucine residues. This dimerization creates a "Y"-shaped structure, where the basic N-terminal regions of each monomer form the arms of the "Y" and fit into the major groove of the DNA on opposite sides of the helix.
- Examples: C/EBP (CAAT/Enhancer-Binding Protein), Fos, Jun, CREB, and c-Myc.
12.4 Zinc Finger Motifs
Zinc finger motifs utilize coordinated zinc ions (Zn2+) to fold a small polypeptide chain into a stable, compact DNA-binding domain. There are two major classes:
12.4.1 Cys2His2 (C2H2) Zinc Finger
- Structure: The most common DNA-binding motif in eukaryotic genomes. Each finger consists of approximately 30 amino acids. A single zinc ion is coordinated tetrahedrally by two conserved cysteine residues on one side of the domain and two conserved histidine residues on the other.
- The consensus sequence of a single finger is: -Cys-X2-4-Cys-X12-His-X2-8-His-
- Mechanism: The tetrahedral coordination folds the peptide into a compact structure consisting of a two-stranded β-sheet packed against an α-helix. The α-helix acts as the recognition domain, fitting into the major groove of the DNA to contact three consecutive base pairs. These fingers are typically arranged in tandem repeats (e.g., three or more fingers in a row), allowing the protein to wrap around the major groove of the DNA double helix.
- Examples: TFIIIA (contains 9 tandem zinc fingers) and Sp1.
12.4.2 Cys4 (C4 / Cys2Cys2) Zinc Finger
- Structure: Found exclusively in the steroid hormone receptor superfamily of nuclear receptors.
- The consensus sequence of a single finger is: -Cys-X2-Cys-X13-Cys-X2-Cys-
- Unlike the C2H2 finger, the C4 zinc finger domain consists of a larger, integrated structure of 70 to 80 amino acids that coordinates two zinc ions, each bound tetrahedrally by four cysteine residues.
- Mechanism: The coordination of the two zinc ions stabilizes a structure that forms two perpendicular α-helices. One helix acts as the recognition helix to bind the major groove of the DNA, while the other mediates homodimerization of the receptor.
- Examples: Estrogen Receptor and Glucocorticoid Receptor.
Chapter 13: Solved Biochemical and Quantitative Problems
Problem 13.1: Promoter Competition and Relative Strengths
You are comparing two different promoters of E. coli genes.
- Promoter A contains the -10 sequence 5′-TATGAT-3′.
- Promoter B contains the -10 sequence 5′-CATGAT-3′.
The wild-type consensus -10 sequence recognized by the primary σ70 subunit is 5′-TATAAT-3′.
Detailed Solution
Consensus Alignment:
- σ70 Consensus: T A T A A T
- Promoter A: T A T G A T (1 mismatch at position 4: A → G)
- Promoter B: C A T G A T (2 mismatches at position 1 [T → C] and position 4 [A → G])
- Biochemical Binding Affinity: The σ70 subunit recognizes the promoter via specific hydrogen bonding and hydrophobic interactions between its Region 2.4 α-helix and the bases of the -10 box.
- DNA Melting Kinetics: The transition from the closed binary complex to the open binary complex requires the melting (unwinding) of the DNA double helix starting at the -10 box. The substitution of a highly conserved thymine (T) to a cytosine (C) at position 1 in Promoter B increases the GC content and introduces a mismatch that significantly reduces the binding affinity of the σ70 subunit.
- Conclusion: Promoter A will be transcribed much more efficiently than Promoter B because its sequence matches the consensus sequence more closely. In transcription, promoters whose sequences are closer to the consensus are stronger because they form the closed binary complex with higher affinity and transition to the open complex at a faster rate.
Problem 13.2: Synthesis Kinetics of a Large Eukaryotic Gene
The human Dystrophin gene is approximately 2.4 × 106 base pairs (2.4 Mb) in length. Eukaryotic RNA Polymerase II transcribes at an average rate of 40 nucleotides per second under physiological conditions. Assume the polymerase transcribes continuously without arresting or falling off.
Detailed Solution
Given Parameters:
- Gene Length (L) = 2,400,000 base pairs
- Transcription Rate (R) = 40 nt/sec
Calculate Time in Seconds (Tsec):
Tsec = L / R = 2,400,000 nt / 40 nt/sec = 60,000 seconds
Convert to Minutes:
Tmin = 60,000 seconds / 60 seconds/minute = 1,000 minutes
Convert to Hours:
Thours = 1,000 minutes / 60 minutes/hour ≈ 16.67 hours
Conclusion: It will take approximately 16.67 hours (or 16 hours and 40 minutes) for a single RNA Polymerase II molecule to transcribe the Dystrophin gene. This long synthesis time highlights why co-transcriptional RNA processing (capping, splicing) occurs simultaneously during transcription elongation, rather than waiting for transcription to finish.
Problem 13.3: Error Rates and Biological Impact
Detailed Solution
- Lack of Genomic Permanence: An error made during DNA replication becomes a permanent mutation in the genome, which is passed on to all future daughter cells and can lead to cell death or cancer. In contrast, an error made during transcription only affects a single molecule of mRNA.
- Turnover and Abundance: mRNA molecules are transient and are rapidly degraded by cellular nucleases (high turnover rate). A single gene is typically transcribed hundreds of times, producing many correct mRNA transcripts. The few defective mRNA molecules containing transcriptional errors will produce a small number of mutated proteins, which are quickly degraded by the proteasome, while the vast majority of correct transcripts will produce normal proteins.
- Lack of Dedicated Proofreading Exonuclease: RNA polymerases lack a 3′ → 5′ editing exonuclease domain equivalent to the one found in DNA polymerases, allowing them to transcribe at a high rate (40 nt/sec) without the thermodynamic and kinetic costs of high-fidelity proofreading.
Problem 13.4: Biochemical Matching Matrix
Match the transcriptional inhibitors with their specific biochemical mechanisms and target enzymes:
| Inhibitors | Mechanisms |
|---|---|
| 1. Rifampicin | A. Intercalates between GC base pairs in DNA, blocking elongation of both prokaryotic and eukaryotic polymerases. |
| 2. α-Amanitin | B. Lacks a 3′-OH group; is incorporated into the growing RNA chain and causes premature termination. |
| 3. Actinomycin D | C. Binds to the eubacterial RNA polymerase β subunit, sterically blocking the exit of the growing RNA chain once it reaches 2-3 nt. |
| 4. Cordycepin | D. Binds to the active site of eukaryotic RNA Polymerase II, physically blocking the translocation step. |
Detailed Solution
- Rifampicin → C: Rifampicin is a eubacterial-specific antibiotic that binds to the β subunit (rpoB) of RNA polymerase. It acts as a steric blocker, preventing the nascent RNA chain from growing past 2-3 nucleotides, thereby blocking promoter clearance.
- α-Amanitin → D: This cyclic peptide binds specifically to eukaryotic RNA Polymerase II (and RNA Polymerase III at much higher concentrations), blocking its translocation along the DNA.
- Actinomycin D → A: This is a non-specific DNA-intercalating drug that inserts its planar ring between adjacent GC base pairs, distorting the template and blocking both prokaryotic and eukaryotic RNA polymerases.
- Cordycepin → B: Cordycepin (3′-deoxyadenosine) lacks a 3′-OH group. Once phosphorylated to cordycepin triphosphate and incorporated into RNA, it prevents the formation of the next phosphodiester bond, resulting in chain termination.
In this lesson
LessonStep 9 of 33

