Gene Expression Analysis, Mutagenesis, and Protein-Macromolecule Interactions

Introduction to Gene Expression Analysis & Sheet-Based Blotting

1. Introduction to Gene Expression Analysis & Sheet-Based Blotting

Northern Blotting · Membrane Hybridization · Dot & Slot Blot Assays

1.1 Biophysical Basis of Hybridization

Gene expression analysis measures how a cell's genome is transcribed into functional RNA molecules. The foundation of these detection techniques lies in nucleic acid hybridization — the highly specific, non-covalent association of complementary single-stranded DNA or RNA molecules through Watson–Crick base pairing (A=T/U and G≡C).

The stability of a hybrid duplex is governed by hydrogen bonding and base-stacking thermodynamics. To identify a specific target transcript within a complex mixture of total cellular RNA, a labeled single-stranded probe — either radioactive 32P or a non-isotopic biotin/digoxigenin-labeled probe — is introduced.

Stringency control: by adjusting environmental conditions — specifically temperature, ionic strength, and denaturants such as formamide — researchers control the stringency of hybridization. High-stringency conditions ensure that only perfectly matched sequences form stable duplexes, preventing non-specific background binding.

1.2 Northern Blotting: Experimental Setup, Membranes, and Run Dynamics

Northern blotting, first developed in 1977 by James Alwine, David Kemp, and George Stark, is a sheet-based hybridization method used to estimate the abundance, size, and splicing variations of specific mRNA transcripts. Unlike Southern blotting, which analyzes DNA, Northern blotting targets RNA.

1.2.1 Denaturing Gel Electrophoresis

Because RNA is single-stranded, it can fold into complex intramolecular secondary structures — hairpins, stem-loops, and pseudoknots — through local base-pairing. These structures alter the hydrodynamic volume and shape of individual RNA molecules, meaning they no longer migrate strictly by molecular weight on a native gel.

To overcome this, RNA is resolved through an agarose gel containing denaturants — most commonly formaldehyde or glyoxal/dimethyl sulfoxide (DMSO).

Mechanism: formaldehyde continuously denatures RNA during electrophoresis by forming reversible adducts with amino groups on the bases, preventing intramolecular hydrogen bonding. This keeps every RNA molecule in an extended, random-coil conformation and migrating through the gel pores at a rate inversely proportional to the log of its molecular weight.

1.2.2 Capillary Transfer (Blotting)

Because agarose gels are thick, fragile, and prone to rapid solute diffusion, the resolved RNA must be transferred and immobilized onto a durable, sheet-like membrane for hybridization. This transfer is driven by capillary action.

Weight (~0.5 kg) Glass plate Paper tissues (absorbs transfer buffer) 3× Dry filter paper Membrane — RNA immobilizes here Agarose gel (inverted) Filter paper wick (drapes into buffer) Support block Transfer buffer — 20× SSC 3.0 M NaCl, 0.3 M trisodium citrate, pH 7.0 Capillary flow of buffer

Figure: Capillary transfer (blotting) apparatus. A solid support block sits in a tray of high-salt transfer buffer (20× SSC) and is draped with a thick filter-paper wick dipping into the buffer. The inverted agarose gel rests on the wick, topped by the membrane, dry filter paper, absorbent paper tissues, a glass plate, and a light weight for even contact.

Run dynamics: the dry paper tissues draw the high-salt buffer up through the wick, the gel, and the membrane via capillary action. As the buffer flows perpendicular to the gel surface, it elutes the RNA molecules out of the gel matrix and carries them onto the membrane surface, where they become physically trapped in a pattern that faithfully replicates the original gel lanes.

1.2.3 Membrane Materials and Immobilization Chemistry

Three primary types of membranes are used, each relying on a distinct chemical interaction to bind single-stranded nucleic acids.

Membrane materialSurface chemistryBinding mechanismBinding capacityKey properties
NitrocelluloseNitroester–cellulose polymerNon-covalent (hydrophobic and electrostatic interactions)≈ 80–100 µg/cm²Fragile, becomes brittle upon baking; requires baking at 80°C under vacuum to prevent ignition
Neutral nylonUnmodified polyamide (–CONH–) backboneNon-covalent (hydrophobic and hydrogen bonding)≈ 100–150 µg/cm²Highly durable, flexible, tear-resistant; excellent for multiple rounds of stripping and re-probing
Charge-modified (positively charged) nylonPolyamide functionalized with quaternary amine groupsElectrostatic binding (positive amines bind negative phosphates)≈ 400–500 µg/cm²Highest sensitivity; binds both single- and double-stranded nucleic acids with extreme affinity
Activated papers (DBM and DPT)Diazobenzyloxymethyl or diazophenylthioetherCovalent (diazo coupling to guanine/uracil bases)ModerateRequires chemical activation; highly stable covalent linkage

Fixation step: once transfer is complete, the nucleic acids are only loosely bound to nitrocellulose or neutral nylon membranes and must be permanently fixed to prevent them from washing off during subsequent high-temperature hybridization and washing steps.

Nitrocellulose: fixed by vacuum baking at 80°C for 2 hours.
Nylon membranes: fixed by UV crosslinking at 254 nm for ≈ 1–2 minutes (dosage of 120,000 µJ/cm²). The UV light induces covalent bonds between thymine (DNA) or uracil (RNA) bases and the amine/amide groups of the nylon matrix.

1.3 Dot Blot and Slot Blot Assays: Quantifying Non-Fractionated Nucleic Acids

When researchers need to analyze a large number of samples and only require quantitative abundance data — without needing to determine molecular size — gel electrophoresis is omitted entirely. Instead, samples are applied directly to the membrane as dots or slots.

Dot Blot Assay (circular deposition) Circular spots — suited for rapid qualitative screening Slot Blot Assay (rectangular deposition) Rectangular slots — ideal for densitometric quantification

Figure: Dot blot vs. slot blot deposition patterns. Both formats apply liquid sample directly to the membrane, bypassing gel electrophoresis entirely, but differ in the geometry of the deposited sample.

Dot blot assay: liquid samples containing DNA or RNA are spotted directly onto the membrane as small circular spots, using a pipette or a vacuum-assisted manifold. After drying, the nucleic acids are denatured on the membrane with a mild alkaline solution, fixed by baking or UV crosslinking, and hybridized with a labeled probe. It is simple, rapid, and ideal for screening the presence or absence of a gene or virus.

Slot blot assay: the sample is deposited through a manifold containing narrow, rectangular slots rather than circular holes. The slot geometry concentrates the sample into a thin, linear band, which is far superior for densitometric scanning and precise quantification because it minimizes signal dispersion and permits a more accurate comparison of band intensity.

High-Resolution Transcript Mapping

2. High-Resolution Transcript Mapping

S1 Nuclease Mapping · Primer Extension Assay

2.1 Why High-Resolution Mapping Is Needed

While Northern blotting provides an estimate of transcript size, it lacks the single-nucleotide resolution required to map the exact transcription boundaries — the 5′ start site and 3′ termination site — or to identify intron–exon borders. Two classical enzymatic methods are used for high-resolution transcript mapping: S1 Nuclease Mapping and the Primer Extension Assay.

2.2 S1 Nuclease Mapping: Detecting 5′ and 3′ Ends of Transcripts

S1 nuclease is an endonuclease from the fungus Aspergillus oryzae that specifically degrades single-stranded DNA and single-stranded RNA into acid-soluble 5′-nucleotides, while leaving double-stranded DNA or DNA–RNA hybrids completely intact.

2.2.1 5′ End Mapping Protocol

  1. Probe preparation: a double-stranded DNA fragment containing the anticipated 5′ transcription start site is selected. The 5′ end of the antisense (template) strand is radioactively labeled using [γ-32P]ATP and T4 Polynucleotide Kinase (T4 PNK).
  2. Denaturation and hybridization: the labeled dsDNA probe is denatured by heat and hybridized under high-stringency conditions with total cellular RNA. The target mRNA hybridizes specifically to the complementary antisense labeled DNA probe, forming a stable DNA–RNA heteroduplex.
  3. S1 digest: the mixture is treated with S1 nuclease. The enzyme rapidly digests the single-stranded, unhybridized 3′ and 5′ overhangs of the DNA probe, as well as any unhybridized single-stranded RNA. Only the double-stranded DNA–RNA hybrid portion survives this enzymatic attack.
  4. Resolution: the reaction is denatured — releasing the single-stranded, protected labeled DNA fragment from its RNA partner — and resolved on a high-resolution denaturing polyacrylamide–urea sequencing gel adjacent to a sequencing ladder. The size of the surviving labeled DNA band on the autoradiogram corresponds to the exact distance from the 5′ labeled site to the transcription start site (+1).
Step 1 — Labeled probe (before hybridization) 3′ sense (unlabeled) 5′* antisense (labeled) Step 2 — Hybridize with total RNA sense (partly ssDNA) mRNA (3′ → overlapping region) 5′* antisense (hybridized region solid) ↑ TSS (+1) Unpaired region (left, dashed) is digested away in the next step → Step 3 — After S1 nuclease digestion digested away (unpaired probe / promoter region) Protected labeled fragment (surviving DNA–RNA hybrid) Fragment length = distance from 5′ label to TSS (+1) read against an adjacent sequencing ladder on denaturing PAGE

Figure: S1 nuclease 5′ end mapping. The 5′-labeled antisense probe hybridizes to mRNA across the transcribed region; S1 nuclease removes every unpaired single-stranded segment, leaving only the protected DNA–RNA hybrid. The surviving labeled fragment's length, read against a sequencing ladder, gives the exact distance from the label to the transcription start site.

2.3 Primer Extension Assay: Determining Transcription Start Sites (+1)

The primer extension assay is a highly sensitive method used to determine the exact +1 transcription initiation site of a gene and to quantify the rate of transcription in vitro.

  1. Primer annealing: a synthetic, single-stranded oligonucleotide primer (typically 18–24 residues long) is labeled at its 5′ end using [γ-32P]ATP and T4 PNK. The primer is designed to be complementary to a region within the mRNA coding sequence, approximately 50–150 nucleotides downstream of the anticipated 5′ end of the transcript, and is annealed specifically to the complementary mRNA template.
  2. Extension reaction: reverse transcriptase (such as AMV or M-MuLV) is added along with an abundant supply of unlabeled dNTPs. The enzyme binds the 3′-OH end of the annealed primer and synthesizes a complementary DNA (cDNA) strand in the 5′→3′ direction, using the mRNA as template.
  3. Termination: the polymerase continues elongation until it reaches the absolute physical 5′ end of the mRNA template, where it runs out of template and dissociates. The 3′ end of this newly synthesized cDNA strand corresponds exactly to the 5′ terminus of the transcript.
  4. Analysis: the RNA template is removed by alkaline hydrolysis or RNase H digestion. The radiolabeled cDNA product is denatured and resolved on a denaturing polyacrylamide sequencing gel. Running this product alongside a DNA sequencing ladder generated from the same gene template allows the researcher to read the exact +1 nucleotide where transcription initiated.
Step 1 — Anneal labeled primer to mRNA 5′ 3′ mRNA +1 (transcription start) 5′* labeled primer Step 2 — Extend with reverse transcriptase (RT) RT extends toward mRNA 5′ end newly synthesized cDNA (3′ end marks the +1 site) Distance from primer 5′ end to cDNA 3′ end = transcription start site

Figure: Primer extension mechanism. A labeled primer anneals downstream of the anticipated transcription start site; reverse transcriptase extends it toward the mRNA 5′ end until it runs out of template. After denaturing the RNA, the single-stranded cDNA is resolved on sequencing PAGE, and its length locates the +1 nucleotide.

High-Throughput Transcriptomics: DNA Microarrays

3. High-Throughput Transcriptomics: DNA Microarrays

Glass Slide cDNA Microarrays · High-Density Oligonucleotide GeneChips

3.1 From Single Genes to Whole Genomes

While Northern blotting, S1 mapping, and primer extension are limited to analyzing one or a few genes at a time, DNA microarrays enable the simultaneous monitoring of expression levels for thousands of genes — or even entire genomes — on a single solid-state chip.

3.2 Glass Slide cDNA Microarrays: Two-Color Competitive Hybridization

The glass slide cDNA microarray, pioneered by Patrick Brown and colleagues at Stanford University, is based on micro-spotting pre-fabricated DNA elements — either PCR-amplified cDNA fragments or long synthetic oligonucleotides — onto a chemically modified glass surface.

3.2.1 Probe Immobilization

Glass slides are chemically treated (e.g., coated with poly-L-lysine, amino-silanes, or epoxy-silanes) to introduce reactive positive charges or covalent linkers. Double-stranded cDNA fragments are printed onto the slide in an orderly grid of thousands of microscopic spots (each ≈ 100–300 µm in diameter) using computerized robotic pin-prick printers or ink-jet deposition. The spotted DNA is subsequently denatured into single strands by heat or alkali and covalently crosslinked to the glass surface.

3.2.2 Probe Preparation and Competitive Hybridization

In a typical differential gene expression experiment — comparing healthy tissue against a cancerous tumor:

Cell Population A (Healthy) Cell Population B (Affected) extract total RNA extract total RNA Total RNA Total RNA RT with fluorophore RT with fluorophore Label with Cy3 (Green) Label with Cy5 (Red) mix equal amounts Co-hybridize onto Chip wash non-specific targets Laser Scanning & Imaging Scanned array spots Green — healthy high (Cy3 ≫ Cy5) Red — affected high (Cy5 ≫ Cy3) Yellow — equal expression (Cy3 ≈ Cy5) Black — not expressed (no signal) Signal ratio computed per spot: Cy5 / Cy3 (Red / Green)

Figure: Two-color competitive hybridization workflow. RNA from two cell populations is separately reverse-transcribed with Cy3 (green) or Cy5 (red) fluorophores, mixed in equal amounts, and co-hybridized onto the same chip. Dual-laser scanning and Cy5/Cy3 ratio analysis reveal which genes are up-, down-, or equally regulated between the two states.

Cy3 (cyanine 3): green-emitting dye, excitation ≈ 550 nm, emission ≈ 570 nm — used to label the healthy-cell cDNA.
Cy5 (cyanine 5): red-emitting dye, excitation ≈ 649 nm, emission ≈ 670 nm — used to label the cancerous-cell cDNA.

3.2.3 Washing and High-Resolution Laser Scanning

Following hybridization, the slide is washed extensively with sodium dodecyl sulfate (SDS) and SSC buffers to remove non-specific, weakly bound cDNA targets. The dry slide is then placed in a high-resolution dual-channel confocal laser scanner: one laser excites Cy3 and records green fluorescence intensity at each spot, while a second excites Cy5 and records red intensity. Software overlays the two images and computes the Cy5/Cy3 signal ratio for every spot.

Spot appearanceRatio patternInterpretation
Bright redCy5 ≫ Cy3Gene is highly upregulated in the cancerous state
Bright greenCy3 ≫ Cy5Gene is highly downregulated in the cancerous state (high expression in healthy cells only)
YellowCy3 ≈ Cy5Gene is expressed at identical, equal levels in both states
BlackNo signalGene is not transcribed or expressed in either cell population

3.3 High-Density Oligonucleotide GeneChips: In Situ Photolithography

High-density oligonucleotide microarrays, developed by Stephen Fodor and commercialized by Affymetrix as GeneChips, are manufactured using a highly sophisticated chemical synthesis method that builds short oligonucleotides directly onto a solid support in situ.

3.3.1 In Situ Photolithography Synthesis

Rather than spotting pre-synthesized DNA, GeneChips use photolithography — a technology borrowed from the semiconductor industry — to synthesize oligonucleotides directly on a quartz substrate.

Step 1 — Quartz coated with photolabile protecting groups (X) O–X O–X O–X O–X Step 2 — UV light through a photolithographic mask deprotects selected columns MASK (blocks) MASK (blocks) UV UV O–X O–H O–X O–H Step 3 — Flood with activated phosphoramidite base (e.g. Adenine) O–X O–A–X O–X O–A–X Adenine couples only to the deprotected O–H sites; blocked columns remain unchangedStep 4 — Repeat with a new mask and a different base (e.g. Guanine) O–G–X O–A–X O–X O–A–X 4 × N synthesis cycles build an N-mer array — over 1,000,000 unique probe coordinates fit within an area smaller than 2 cm² (schematic: 4 of many thousands of synthesis columns shown)

Figure: In situ photolithographic synthesis. A photolithographic mask selectively deprotects chosen coordinates with UV light; an activated, protected phosphoramidite base couples only at the deprotected sites. Repeating this cycle with a new mask and a different base builds up a unique oligonucleotide sequence at every coordinate on the chip.

3.3.2 Probe Pair Design (PM vs. MM)

To ensure absolute specificity and eliminate false positives caused by non-specific cross-hybridization, Affymetrix GeneChips do not rely on a single probe sequence per gene. Instead, they use Probe Pairs:

Perfect Match (PM) — 25-mer exactly complementary to the target
Mismatch (MM) — identical 25-mer, but the 13th (middle) base is swapped for its Watson–Crick complement
Analysis: because the mismatch probe differs by only a single base, non-specific hybridization or background binding affects both the PM and MM probes equally. The true, specific hybridization signal is calculated as Signal = IPM − IMM.
Protein-Protein Interaction Assays

4. Protein-Protein Interaction Assays

Phage Display · Yeast Two-Hybrid · Y1H / Y3H / Reverse Y2H

4.1 Mapping Physical Interactions Between Proteins

Understanding biological pathways requires mapping physical interactions between proteins. Two classical methods used to discover and characterize these interactions are Phage Display and the Yeast Two-Hybrid (Y2H) Assay.

4.2 Phage Display: M13 Biology, Fusion Design, and Biopanning

Phage display, invented by George P. Smith in 1985, is an in vitro screening technique that physically links a protein's genotype (the DNA inside the phage) to its phenotype (the protein displayed on the outside).

4.2.1 M13 Filamentous Phage Biology

M13 is a non-lytic, filamentous bacteriophage that infects Escherichia coli carrying the F-plasmid. It contains a circular, single-stranded DNA genome wrapped in a protein envelope.

pVIII — major coat protein (≈2,700 copies along the cylinder) pVII / pIX (5 copies each) pIII / pVI (5 copies each) Fusion Protein displayed on pIII Single-stranded circular DNA genome (packaged inside the coat)

Figure: M13 filamentous phage structure. The cylindrical body is built from thousands of copies of the major coat protein pVIII; the two tips each carry five copies of a pair of minor coat proteins. A protein or peptide fused to pIII is displayed at one tip while the encoding DNA remains packaged inside.

pIII fusions: only 5 copies per phage — suited to displaying larger, folded proteins or diverse peptide libraries for high-affinity monovalent binding screens.
pVIII fusions: ≈2,700 copies per phage — suited to short peptides (typically <10 residues) for high-density, multivalent binding displays, though steric hindrance limits insertion of larger proteins.

4.2.2 Phagemid Vector Systems

To facilitate genetic engineering, researchers use phagemids — hybrid plasmid vectors containing both a standard plasmid origin of replication (ColE1) and an M13 filamentous phage origin of replication. The phagemid lacks the genes required for phage assembly and packaging.

Helper phage: to produce display particles, E. coli cells carrying the phagemid are co-infected with a helper phage (e.g., M13K07). The helper phage supplies all the structural proteins (wild-type pIII and pVIII), while the phagemid's M13 origin directs packaging of the recombinant phagemid DNA into the newly assembled virions — producing a "display library" where each phage displays a unique fusion protein on its surface and carries the corresponding gene inside its capsule.

4.2.3 The Biopanning Cycle

To isolate a specific high-affinity binder against a target molecule from a library of 109 diverse phages, a multi-step selection procedure called biopanning is performed.

1. Incubate phage library with target immobilized on a solid support 2. Wash aggressively to remove non-specific / weakly bound phages 3. Elute high-affinity bound phages low-pH glycine buffer or free ligand 4. Infect E. coli and amplify recovered phages to repeat the cycle Repeat for 3–4 rounds of increasing stringency Sequence phagemids from isolated clones

Figure: The biopanning receptor cycle. The phage library is incubated with immobilized target, washed to remove weak binders, and the tightest binders are eluted and amplified by re-infecting E. coli. Repeating the cycle for 3–4 rounds with progressively harsher washing enriches for the highest-affinity binders, which are finally sequenced.

4.3 Yeast Two-Hybrid (Y2H) Assay: Classical GAL4 and LexA Reconstitution

The Yeast Two-Hybrid (Y2H) assay, developed by Stanley Fields and Ok-Kyu Song in 1989, is an in vivo genetic system used to detect protein-protein interactions inside the nucleus of the budding yeast Saccharomyces cerevisiae.

4.3.1 Reconstitution of Transcription Factors

Y2H is based on the modular nature of eukaryotic transcription factors (such as GAL4 or the bacterial repressor LexA), which contain two physically distinct, autonomous domains: a DNA-Binding Domain (DBD), which binds a specific upstream activating sequence (UAS) on the promoter, and an Activation Domain (AD), which recruits RNA Polymerase II. If separated, the DBD can bind DNA but cannot initiate transcription, and the AD remains functional but cannot locate the promoter — but if brought into physical proximity, their reconstitution drives transcription of a reporter gene.

BAIT (Protein X) PREY (Protein Y) X Y Protein-Protein Interaction DBD AD UAS Site AD is tethered to the promoter via the bait–prey interaction Promoter region transcription initiates Reporter Gene (lacZ / HIS3) — transcribed

Figure: Y2H molecular mechanism. Bait (fused to the DBD) anchors at the UAS on the promoter; if bait and prey physically interact, the prey's AD is brought into proximity with the promoter, reconstituting a functional transcription factor and driving reporter gene expression.

4.3.2 Experimental Setup

Two recombinant fusion plasmids are constructed and transformed into a single yeast strain: the bait plasmid fuses the protein of interest to the DNA-binding domain (usually GAL4-DBD or LexA), and the prey plasmid fuses a target partner protein or a cDNA library to the activation domain (usually GAL4-AD or B42 acid blob). The host strain carries integrated reporter genes — such as lacZ (β-galactosidase), HIS3 (histidine biosynthesis), or ADE2 (adenine biosynthesis) — under the control of the promoter recognized by the DBD.

4.3.3 Interaction Outcomes

OutcomeReporter statusGrowth / color phenotype
No interactionAD stays separated from the promoter; reporter is silentCannot grow on histidine/adenine-deficient media; colonies appear white with X-gal
Positive interactionAD is tethered to the DBD at the promoter; RNA Pol II is recruitedGrows on selective media; turns bright blue with X-gal (β-galactosidase active)

4.4 Y1H, Y3H, and Reverse Two-Hybrid (rY2H) Variations

The versatility of the two-hybrid principle has led to several specialized variations.

VariantPurposeMechanism
Yeast One-Hybrid (Y1H)Identify protein–DNA interactions (e.g., which transcription factors bind a promoter/enhancer)Target DNA is integrated upstream of a reporter; a library of GAL4-AD fusions is introduced — a fusion that binds the DNA element directly recruits the AD and activates the reporter. No DBD fusion is required.
Yeast Three-Hybrid (Y3H)Study RNA–protein or ligand-mediated interactions via a third componentA hybrid "RNA bridge" has one loop binding a DBD-fused RNA-binding protein (e.g., MS2 coat protein) and a second loop binding the target RNA. A library of AD fusions is introduced; binding to the target RNA loop reconstitutes the transcription factor.
Reverse Two-Hybrid (rY2H)Identify mutations, drugs, or peptides that disrupt a known interactionA counter-selectable reporter (URA3) converts 5-FOA into a toxic compound. If bait/prey interact, URA3 is expressed and cells die on 5-FOA; an inhibitor or disrupting mutation silences URA3, letting cells survive — a high-throughput screen for drug discovery.
Protein-DNA / Macromolecule Interactions

5. Protein-DNA / Macromolecule Interactions

EMSA · DNase I Footprinting · DMS Footprinting · Modification Interference

5.1 Genomic Regulation and Protein–DNA Interactions

Genomic regulation relies on proteins interacting specifically with DNA. Four classical methods are used to identify and map these interactions: EMSA, DNase I Footprinting, DMS Footprinting, and Modification Interference Assays.

5.2 Electrophoretic Mobility Shift Assay (EMSA / Gel Retardation)

The Electrophoretic Mobility Shift Assay (EMSA), also known as the gel shift or band shift assay, is a rapid and highly sensitive method used to determine if a protein has binding affinity for a specific double-stranded DNA or RNA fragment.

5.2.1 Biophysical Principles

EMSA is based on the physical behavior of molecules migrating through a non-denaturing gel matrix under an electric field. Small, double-stranded DNA fragments (typically 100–300 bp) carry a uniform, high negative charge density and migrate rapidly toward the anode. If a protein binds specifically to the DNA, the resulting complex has a much larger hydrodynamic size, an altered shape, and a reduced net negative charge (basic residues in the DNA-binding domain neutralize some phosphates). When resolved on a native gel, this bulky complex is physically retarded, migrating much slower than free DNA — producing a distinct, upward "shift" of the labeled band.

+ Migration toward the anodeLane 1: Free DNA Probe Lane 2: DNA + Binding Protein [ Well ] [ Well ] Free DNA (unbound) Shifted complex (Protein–DNA) ↑ where free DNA would run (faint — most probe is now shifted)

Figure: EMSA native gel resolution. Free DNA migrates quickly to a low position in the gel. When a protein binds the probe, the larger, less negatively charged complex is retarded and appears as a distinct upward-shifted band closer to the well.

Reagent rules: native PAGE gels are used because denaturing agents (SDS, urea, or heat) would instantly disrupt the non-covalent hydrophobic and electrostatic linkages holding the protein–DNA complex together.
Super-shift assay: to confirm the identity of the binding protein in a complex cellular extract, an antibody specific to the suspected protein is added. The antibody binds the complex, forming an even larger "antibody–protein–DNA" ternary complex that migrates even more slowly — an ultra-slow "supershift" band higher on the gel.

5.3 DNase I Footprinting (Nuclease Protection Footprinting)

While EMSA indicates if a protein binds DNA, DNase I Footprinting identifies the exact sequence of nucleotides protected by that protein.

  1. Probe labeling: a double-stranded DNA fragment containing the target binding region is labeled at the 5′ end of only one strand using [γ-32P]ATP and T4 PNK.
  2. Protein binding: the labeled DNA is split into Tube A (free, naked DNA control) and Tube B (incubated with the purified DNA-binding protein).
  3. Limited digestion: both tubes undergo limited digestion with DNase I, an endonuclease with little sequence specificity.
  4. The "single-hit" rule: reaction conditions are optimized so that, on average, each DNA molecule is cleaved only once along its length, generating a complete population of fragments of every possible length.
  5. Protection: in Tube B, the bound protein physically blocks DNase I from cleaving phosphodiester bonds within its binding site.
  6. Resolution: the reactions are denatured and resolved side-by-side on a high-resolution denaturing polyacrylamide–urea sequencing gel next to a DNA sequencing ladder.
Lane A: Free DNA Lane B: DNA + Protein ← The footprint (region protected by the protein)

Figure: DNase I footprinting gel comparison. Lane A (free DNA) shows a continuous ladder of bands at every cleavable position. Lane B (DNA + protein) shows a gap where bands are absent — the "footprint" — corresponding to the exact region the protein occupied. Aligning this gap with an adjacent sequencing ladder reveals the exact protected sequence.

5.4 Modification Protection Footprinting (Dimethyl Sulfate & Piperidine Cleavage)

Because DNase I is a large protein (31 kDa), it experiences steric hindrance, meaning the footprint it generates is often larger than the actual sequence contacted by the protein. To achieve higher resolution, researchers use small chemical cleavage agents in Modification Protection Footprinting, most commonly Dimethyl Sulfate (DMS).

  1. Methylation step: end-labeled DNA is incubated with or without the DNA-binding protein. DMS — a small methylating agent that freely penetrates the DNA–protein complex — is added, and specifically methylates the N7 position of guanine residues in the major groove.
  2. Protection: guanines directly contacted by the binding protein are physically shielded from DMS attack; guanines in unprotected regions are methylated.
  3. Cleavage step: the protein is removed, and the DNA is treated with piperidine, which depurinates methylated guanines and catalyzes a β-elimination reaction that breaks the sugar–phosphate chain at those positions.
  4. Resolution: fragments are resolved on a denaturing polyacrylamide gel. Protected guanosines are not methylated and therefore not cleaved — appearing as empty gaps in the guanine sequencing ladder, providing precise, single-nucleotide contact resolution.
FeatureDNase I footprintingDMS / piperidine footprinting
Cleavage agentEndonuclease (31 kDa protein)Small chemical (DMS) + piperidine cleavage
ResolutionLower — footprint often larger than the true contact region (steric hindrance)Higher — single-nucleotide contact resolution
Base specificityLittle sequence specificityMethylates guanine N7 positions specifically

5.5 Modification Interference Assays

While footprinting assays show which nucleotides are protected by a bound protein, Modification Interference Assays identify which specific functional groups on the DNA are essential for protein binding to occur.

Global chemical modification one modified base per molecule (DMS or formic acid) Incubate with DNA-binding protein Resolve by EMSA (gel shift) Bound fraction (shifted) still capable of binding the protein Free fraction (unbound) rejected — modification disrupted binding Cleave with piperidine Cleave with piperidine Resolve both fractions on a sequencing gel bands absent only in the bound lane = essential contacts

Figure: Modification interference workflow. DNA carrying one random modified base per molecule is incubated with the protein and separated by EMSA into bound and free fractions. Both are cleaved at the modified sites with piperidine and resolved side-by-side; a modified position present in the free fraction but missing from the bound fraction marks an essential contact point.

Outcome: modified bands present in the free fraction but completely absent from the bound fraction represent nucleotides where modification directly interfered with protein binding — these residues are essential contact points required for recognition.
In Vitro Site-Directed Mutagenesis

6. In Vitro Site-Directed Mutagenesis

Cassette & M13 Primer Extension · Overlap Extension, Megaprimer & Inverse PCR · Random Mutagenesis

6.1 Making Precise, Targeted Changes to DNA

Deciphering protein function requires making precise, targeted changes to the encoding DNA sequence. Site-directed mutagenesis is an in vitro genetic engineering technique used to create specific insertions, deletions, or nucleotide substitutions. These methods are divided into Non-PCR-Based and PCR-Based strategies.

6.2 Non-PCR-Based Site-Directed Mutagenesis

6.2.1 Cassette Mutagenesis

Cassette mutagenesis is a classical method used to introduce large-scale substitutions or multiple mutations within a localized region of a gene.

Step 1 — Identify flanking restriction sites (RE1, RE2) RE1 Wild-Type Gene RE2Step 2 — Cleave with RE1/RE2; discard the wild-type fragment RE1 (fragment removed) RE2Step 3 — Ligate a synthetic mutant cassette in its place RE1 Mutant Cassette (*) RE2 Recombinant plasmid carrying the desired mutation(s)

Figure: Cassette mutagenesis. The gene is cloned with two unique flanking restriction sites; the plasmid is cut and the wild-type fragment discarded, then a synthetic double-stranded oligonucleotide carrying the desired mutations is ligated in its place.

Evaluation: simple, highly efficient, and can achieve multiple, complex codon changes simultaneously. The major limitation is that it strictly requires unique, flanking restriction sites in close proximity to the mutation target.

6.2.2 Primer Extension Mutagenesis (M13 Heteroduplex Segregation)

This method, pioneered by Michael Smith (1993 Nobel Prize in Chemistry), utilizes single-stranded M13 bacteriophage DNA templates.

  1. The target gene is cloned into the double-stranded replicative form (RF) of M13, and single-stranded viral DNA (the (+) strand) is harvested from secreted phage particles.
  2. A synthetic oligonucleotide primer (≈20 residues) is synthesized with the desired single base mismatch in the middle, flanked by perfectly complementary sequences on either side, and annealed to the M13 template.
  3. Klenow Fragment (DNA Polymerase I lacking 5′→3′ exonuclease activity, to protect the mutagenic primer) extends the primer around the circular template; T4 DNA Ligase seals the nick, forming a closed circular heteroduplex — one strand wild-type, the complementary strand mutant.
  4. The heteroduplex is transformed into E. coli. Semi-conservative replication segregates the mutant and wild-type strands into separate progeny plaques, identified by plaque hybridization with the labeled mutagenic primer under stringent conditions.
Step 1 — Anneal mutagenic primer to ssDNA template 3′ 5′ (ssDNA template) primer (mismatch *)Step 2 — Extend with Klenow + seal with DNA ligase Closed circular heteroduplex: outer strand mutant (*), inner strand wild-typeStep 3 — Transform into E. coli; strands segregate on replication Mutant progeny (*) identified by probe hybridization Wild-type progeny probe dissociates on wash

Figure: M13 primer extension mutagenesis. A mutagenic primer anneals to the single-stranded template and is extended and ligated into a closed heteroduplex. On transformation and replication, the strands segregate into mutant and wild-type progeny, which are distinguished by hybridization stringency with the labeled primer.

6.3 PCR-Based Site-Directed Mutagenesis

With the advent of thermostable DNA polymerases, PCR-based methods have become the standard due to their speed, simplicity, and lack of requirement for single-stranded viral vectors.

6.3.1 Overlap Extension PCR

A highly versatile method used to introduce site-specific point mutations, insertions, or deletions within any region of a gene.

Reaction A: flanking Primer 1 + mutagenic Primer 3 → Product AB 1 3 (*) template Product AB (mutation at right end)Reaction B: mutagenic Primer 2 + flanking Primer 4 → Product CD 2 (*) 4 Product CD (mutation at left end)Mix products, denature, and reanneal at the overlapping mutant site extend overlapping 3′ endsFinal amplification with flanking Primers 1 and 4 Full-length mutant gene product

Figure: Overlap extension PCR. Two parallel reactions each place the mutation at one end of an overlapping product. After mixing, denaturing, and reannealing at the shared mutant overlap, polymerase extends the overlap and flanking primers amplify the full-length mutant gene.

6.3.2 Megaprimer PCR

An efficient, three-primer PCR strategy that avoids the need to purify and combine two separate intermediate PCR products.

First PCR: flanking Primer 1 + mutagenic Primer A → the "megaprimer" 1 A (*) Megaprimer (short mutant dsDNA fragment)Second PCR: megaprimer + flanking Primer 2 → full-length product 2 megaprimer anneals & primes synthesis Taq extends to full length; PCR amplifies

Figure: Megaprimer PCR. A first PCR round with a flanking primer and an internal mutagenic primer generates a short mutant "megaprimer." In a second PCR round, this megaprimer anneals to the template and is extended, and a second flanking primer completes and amplifies the full-length mutant gene.

6.3.3 Inverse PCR Method

A rapid method used to introduce mutations directly into a whole circular plasmid vector without any sub-cloning steps.

Step 1 — Back-to-back primers (P1, P2) flank the mutation site Circular plasmid target gene mutation site P1 → ← P2 Step 2 — PCR amplification around the entire plasmid Linear mutant amplicon (mutation carried at both ends) Step 3 — Circularize (ligate 5′ phosphates, or homologous recombination in vivo) Circularized mutant plasmid Step 4 — Transform into E. coli; treat the reaction with DpnI DpnI digests methylated parental DNA unmethylated in vitro mutant plasmid survives

Figure: Inverse PCR mutagenesis. Back-to-back primers carrying the mutation amplify around the entire circular plasmid, producing a linear amplicon that is recircularized by ligation or in vivo recombination. Treating the transformation with DpnI selectively destroys the methylated parental template, leaving only the unmethylated mutant plasmid intact.

6.4 Random and Extensive Mutagenesis

Sometimes, rather than introducing a specific point mutation, researchers want to generate a diverse library of random mutations throughout a target gene for directed evolution or to identify functional domains. Two primary strategies are used.

StrategyMechanismKey manipulation
Error-prone PCRStandard Taq lacks proofreading; deliberately compromising its fidelity forces frequent misincorporationIncreased Mg2+ and added Mn2+ (reduces nucleotide specificity); unbalanced dNTP ratios; high Taq concentration, many cycles, low annealing temperature
Base analog incorporationDegenerate base analogs (e.g., 8-oxo-dGTP, dPTP) structurally mimic standard bases but pair ambiguouslydPTP pairs with both adenine and guanine with similar efficiency; Taq randomly incorporates analogs, which then template insertion of either complementary base in later cycles — producing transitions and transversions at random sites

In this lesson

Scroll to Top