DNA Libraries, Genetic Markers, and Genome Mapping

DNA Libraries

DNA Libraries

Genomic & cDNA Library Construction · Clarke-Carbon Statistics · Library Screening

1. DNA Libraries

A DNA library is a comprehensive collection of cloned DNA fragments representing either the entire genome of an organism or the subset of genes expressed as messenger RNA (mRNA) in a specific cell type or developmental state. Each recombinant clone in a library contains a unique DNA insert carried within a vector backbone and propagated inside a host bacterial or yeast cell colony.

Genomic DNA Library Construction Target Genomic DNA (High Molecular Weight) Digestion with Restriction Enzyme (e.g., Partial Sau3A1 Cleavage) Cleaved DNA Fragments (Overlapping Insert Pool) Linearized Vector (e.g., λEMBL3 / BAC) In Vitro Ligation Recombinant Vectors Transformation / Packaging Transformed Bacterial Host Cells Selective Agar Culture (Millions of Individual Clones)

Figure: Genomic DNA Library Construction. High molecular weight genomic DNA is partially digested into overlapping fragments and ligated in vitro to a linearized vector. The resulting recombinant vectors are transformed or packaged into host bacterial cells, which are plated on selective agar to generate millions of individually identifiable clones.

1.1 Genomic DNA Libraries

A genomic DNA library contains clones representing the total nuclear and organellar DNA sequence of an organism, encompassing both coding regions (exons) and non-coding sequences (introns, intergenic regions, regulatory promoters, enhancers, and repetitive elements).

1.1.1 Shotgun Cloning Strategy

  • Step 1
    Genomic Fragmentation

    Purified genomic DNA is subjected to partial restriction digestion (typically using a frequent-cutting 4-bp restriction endonuclease like Sau3A1, 5′-^GATC-3′) or mechanical shearing. Partial digestion ensures that cleavage occurs stochastically across only a fraction of target sites, producing a population of overlapping fragments of optimal size for vector insertion.

  • Step 2
    Vector Ligation

    The overlapping fragments are ligated into a suitable cloning vector (such as bacteriophage λ, cosmid, BAC, or YAC) using T4 DNA ligase.

  • Step 3
    Host Transformation & Storage

    Recombinant vectors are introduced into competent host cells (E. coli) via transformation, electroporation, or in vitro bacteriophage packaging, generating millions of independent transformed colonies.

1.1.2 Statistical Representation: The Clarke-Carbon Formula

To ensure that a genomic library contains every unique sequence of an organism's genome with a high degree of statistical confidence, the total number of independent recombinant clones (N) required is calculated using the Clarke-Carbon equation:

N =  ln(1 − P)ln(1 − a/b) The Clarke-Carbon Equation

Where:

  • N

    Number of independent recombinant clones required.

  • P

    Desired fractional probability that any specific genomic sequence is present in the library (e.g., P = 0.99 for 99% representation).

  • a

    Average insert size of the cloned genomic DNA fragments (in base pairs).

  • b

    Total size of the target haploid genome (in base pairs).

Clarke-Carbon Mathematical Relationship Larger Insert Size (a) Larger a/b Ratio Fewer Clones Needed (N) Larger Genome Size (b) Smaller a/b Ratio More Clones Needed (N)

Figure: Clarke-Carbon Relationship. Larger cloned insert sizes increase the a/b ratio and reduce the number of clones needed for full genome representation, while larger target genomes decrease the a/b ratio and demand a correspondingly larger clone library.

Sample Quantitative Problem 1

A recombinant genomic library is constructed in the bacteriophage vector λEMBL3 using 20,000-bp (20 kb) inserts generated from a partial Sau3A1 digest of the human genome (3 × 109 bp). Calculate the minimum number of independent recombinant phage clones (N) required to achieve a 99% probability (P = 0.99) of isolating a specific gene contained completely within a 20,000-bp fragment.

Solution:

  1. Identify the given variables: P = 0.99 ⇒ (1 − P) = 0.01;   a = 20,000 bp = 2 × 104 bp;   b = 3 × 109 bp.
  2. Calculate the fractional genome ratio a/b:
    a/b =  2 × 1043 × 109  = 6.6667 × 10−6
  3. Apply the natural logarithm to numerator and denominator:
    ln(1 − P) = ln(0.01) = −4.60517
    ln(1 − a/b) = ln(1 − 6.6667 × 10−6) ≈ −6.6667 × 10−6
  4. Solve for N:
    N =  −4.60517−6.6667 × 10−6  ≈ 690,775 clones
Answer: N ≈ 690,775 clones (using exact floating-point precision, N ≈ 688,000 clones).

1.2 Complementary DNA (cDNA) Libraries

A cDNA library consists of synthetic double-stranded DNA molecules synthesized from mature mRNA templates isolated from a specific tissue, organ, or physiological state. Because cDNA is derived from post-transcriptionally spliced mRNA, it is devoid of introns and non-transcribed regulatory promoter sequences. Consequently, eukaryotic cDNA can be directly expressed inside bacterial host cells to yield functional recombinant proteins.

cDNA Synthesis Workflow from mRNA 1 5′ CAP AAAAAAAA-3′ Mature, spliced mRNA template 2 5′ CAP AAAAAAAA-3′ TTTTTTTT-5′ Oligo(dT) primer anneals to the poly(A) tail 3 5′ CAP AAAAAAAA-3′ 3′-cDNA TTTTTTTT-5′ Reverse Transcriptase First-strand cDNA synthesized (5′ → 3′) using dNTPs 4 5′ CAP AAAAAAAA-3′ 3′-cDNA TTTTTTTT-5′ RNase H / NaOH RNA template degraded — single-stranded cDNA remains 5 AAAAAAAA-3′ TTTTTTTT-5′ DNA Polymerase I 3′ hairpin loop primes second-strand synthesis (ds-cDNA) 6 AAAAAAAA-3′ TTTTTTTT-5′ S1 Nuclease Loop cleaved → blunt, double-stranded cDNA

Figure: cDNA Synthesis Workflow. An oligo(dT) primer anneals to the poly(A) tail of mature mRNA, priming first-strand cDNA synthesis by reverse transcriptase. The RNA template is then degraded, a transient 3′ hairpin loop primes second-strand synthesis by DNA Polymerase I, and S1 nuclease cleaves the loop to yield blunt-ended, double-stranded cDNA.

1.2.1 Multi-Step Enzymatic Cascade of cDNA Synthesis

1First-Strand Synthesis
  • Primer Annealing: A synthetic oligo(dT) primer (a short oligonucleotide sequence of 12–18 deoxythymidines) is annealed to the polyadenylated 3′-poly(A) tail of mature eukaryotic mRNA.
  • Reverse Transcription: Reverse Transcriptase (RNA-directed DNA polymerase) synthesizes a complementary single-stranded DNA (ss-cDNA) strand in the 5′ → 3′ direction using deoxynucleotide triphosphates (dNTPs).
2Template RNA Removal

The RNA strand of the resulting RNA-cDNA heteroduplex is partially or completely digested using RNase H (which specifically degrades RNA in RNA-DNA hybrids) or alkali hydrolysis (0.1 M NaOH).

3Second-Strand Synthesis
  • The 3′ end of the single-stranded cDNA frequently folds back on itself to form a transient hairpin loop structure through self-complementary base pairing.
  • DNA Polymerase I uses this 3′ hairpin loop as a primer to synthesize the second (complementary) DNA strand (5′ → 3′), completing double-stranded cDNA (ds-cDNA).
4Hairpin Loop Cleavage

The single-stranded hairpin loop connecting the two strands is cleaved by S1 Nuclease (a single-strand-specific nuclease), yielding a fully duplex, blunt-ended ds-cDNA molecule ready for adaptor ligation and vector cloning.

1.2.2 Comparative Matrix: Genomic Library vs. cDNA Library

AttributeGenomic DNA LibrarycDNA Library
Starting MaterialTotal chromosomal DNA purified from nucleusMature polyadenylated mRNA purified from cytosol
Sequence RepresentationComplete genome (exons, introns, promoters, repeats)Coding sequences (exons) only; no introns or promoters
Cell/Tissue UniformityIdentical across all cell types of an organismVaries drastically depending on cell/tissue expression
Clone AbundanceEqual representation of all non-repetitive genesProportional to steady-state mRNA abundance levels
Prokaryotic ExpressionIncompatible (bacteria cannot splice eukaryotic introns)Compatible (direct expression into functional proteins)
Average Insert SizeVariable (15 kb to 500 kb depending on vector)Smaller (0.5 kb to 8 kb, representing protein ORFs)

1.3 High-Efficiency Library Screening

Once a DNA library is constructed, specific target clones are identified from millions of background colonies or phage plaques using nucleic acid hybridization or functional expression assays.

Colony & Plaque Hybridization Procedure Agar Culture Plate (Bacterial Colonies / Plaques) Nitrocellulose Overlay Nitrocellulose Membrane Replica (Transferred Cells / Plaques) Cell Lysis & DNA Denaturation (0.5 M NaOH) Immobilized ssDNA Bound to Membrane Hybridization with Labeled Single-Stranded Probe Annealed Target-Probe Duplexes (Unbound Probe Washed Away) Autoradiography / Chemiluminescence Dark Film Signal (Identifies Target Clone on Master Plate) Align to Original Master Plate Positive Clone Isolated from Master Plate for Expansion

Figure: Colony & Plaque Hybridization. A membrane replica of an agar culture is lysed and denatured to expose single-stranded DNA, which is hybridized to a labeled probe. After washing away unbound probe, autoradiography reveals dark signal spots that are aligned back to the master plate to isolate the corresponding positive clone.

1.3.1 Colony and Plaque Hybridization

1Membrane Transfer

A sterile nitrocellulose or nylon membrane disk is pressed onto the surface of an agar master plate containing bacterial colonies or viral plaques, creating an exact replica pattern of immobilized cells.

2In Situ Cell Lysis & DNA Denaturation

The membrane is treated with alkaline detergent solutions (0.5 M NaOH / SDS) to lyse bacterial cell walls, release genomic/plasmid DNA, and denature double-stranded DNA into single strands.

3DNA Immobilization

The single-stranded DNA is permanently crosslinked to the membrane matrix by baking at 80°C under vacuum or exposing to ultraviolet (UV) radiation.

4Probe Hybridization

The membrane is incubated in a hybridization buffer containing a radiolabeled (³²P) or fluorophore-tagged single-stranded oligonucleotide or cDNA probe. The probe selectively anneals to complementary sequences on the membrane.

5Detection & Alignment

Unbound probe is washed away under controlled stringency conditions (salt concentration and temperature). The membrane is exposed to X-ray film (autoradiography). Dark spots on the film correspond to positive recombinant clones, which are isolated from the original master plate.

1.3.2 Synthetic Oligonucleotide Probe Design & Degeneracy

When screening a cDNA library for a gene whose protein sequence is partially known, synthetic oligonucleotide probes are reverse-translated from the amino acid sequence. Because the genetic code is degenerate (redundant), multiple codon combinations can code for the same amino acid. To minimize the pool complexity of synthetic probes, target regions containing low-degeneracy amino acids (such as Methionine and Tryptophan, encoded by 1 codon each) are prioritized.

Degeneracy in Probe Pool Complexity Peptide P: Met (1) × Leu (6) × Arg (6) × Leu (6) = 216 216 unique oligonucleotides — large, non-specific probe pool Peptide Q: Met (1) × Trp (1) × Cys (2) × Trp (1) = 2 2 unique oligonucleotides — small, highly specific probe pool Peptide Q is Preferred Lower degeneracy reduces background noise

Figure: Degeneracy in Probe Pool Complexity. Peptide regions dominated by high-degeneracy amino acids (Leucine, Arginine — 6 codons each) generate very large probe pools, while regions dominated by low-degeneracy amino acids (Methionine, Tryptophan — 1 codon each) generate small, highly specific probe pools that minimize non-specific hybridization.

Sample Quantitative Problem 2

You wish to clone a gene encoding a specific metabolic enzyme. Partial peptide sequencing reveals two candidate four-amino-acid regions: Peptide P: Met-Leu-Arg-Leu; Peptide Q: Met-Trp-Cys-Trp. Determine which peptide sequence would be most suitable for designing a synthetic oligonucleotide probe pool to screen a cDNA library.

Solution:

  1. Analyze codon redundancy for each amino acid: Met = 1 codon (5′-AUG-3′); Trp = 1 codon (5′-UGG-3′); Cys = 2 codons (5′-UGU-3′, 5′-UGC-3′); Leu = 6 codons; Arg = 6 codons.
  2. Calculate total probe pool degeneracy (D) for Peptide P:
    DP = 1 (Met) × 6 (Leu) × 6 (Arg) × 6 (Leu) = 216 unique oligonucleotides
  3. Calculate total probe pool degeneracy (D) for Peptide Q:
    DQ = 1 (Met) × 1 (Trp) × 2 (Cys) × 1 (Trp) = 2 unique oligonucleotides
Conclusion: Peptide Q is significantly better because its degeneracy pool contains only 2 distinct oligonucleotide sequences. A smaller pool size reduces non-specific background hybridization during library screening.
Genetic Markers

Genetic Markers

Classical & Molecular DNA Markers · RFLP Southern Blotting · RAPD, AFLP, SSR & SNP Systems

2. Genetic Markers

A genetic marker is any biological feature, gene sequence, or identifiable chromosomal locus that distinguishes individuals, strains, or species from one another. Genetic markers serve as physical landmarks to trace inheritance, construct genetic linkage maps, and locate traits of agronomic or medical importance.

Classification of Genetic Markers Genetic Markers Classical Markers Molecular DNA Markers Morphological Markers Cytological Markers Biochemical / Protein (Isozymes & Allozymes) PCR-Based Hybridization-Based (RFLP) Arbitrary Primed (RAPD, AFLP) Sequence Specific (SSLP/SSR, SNP)

Figure: Classification of Genetic Markers. Genetic markers are broadly divided into classical markers (morphological, cytological, and biochemical/protein-based) and molecular DNA markers, which are further split into hybridization-based systems (RFLP) and PCR-based systems (arbitrary-primed and sequence-specific).

2.1 Classical Markers

  • Type
    Morphological Markers

    Visually observable phenotypic traits (e.g., flower color, seed coat texture, plant height). Limitations: highly subject to environmental influences, epistatic interactions, and developmental stage; limited in total number.

  • Type
    Cytological Markers

    Structural features of chromosomes observed via microscopic staining (e.g., G-banding patterns, secondary constrictions, translocation breakpoints, heterochromatic knobs).

Biochemical (Protein) Markers

  • Protein
    Allozymes

    Variant forms of a specific enzyme encoded by different alleles at the exact same locus. Allozymes exhibit minor amino acid substitutions that alter net electrical charge, allowing separation via gel electrophoresis.

  • Protein
    Isozymes

    Enzymes that perform the same catalytic function but are encoded by different non-allelic genes located at distinct genomic loci.

2.2 Molecular DNA Markers

DNA markers identify physical sequence variations directly at the chromosomal level. They are non-phenotypic, environmentally neutral, highly abundant across the genome, and reproducible.

Criteria for an Ideal DNA Marker

  • High Polymorphism

    Must exhibit high intra-species or inter-species allelic variability.

  • Even Genomic Distribution

    Must be uniformly scattered across all chromosomes (exons, introns, intergenic space).

  • Co-Dominant Expression

    Must discriminate between homozygous individuals (AA or aa) and heterozygous individuals (Aa).

  • Reproducibility & Stability

    Must yield consistent results across laboratories.

Co-Dominant vs. Dominant Marker Gel Patterns Co-Dominant Marker (Distinguishes Heterozygotes)P1 (AA) P2 (aa) F1 (Aa) 3 Distinct Band Profiles Heterozygote Clearly Identified Dominant Marker (Binary Presence / Absence)P1 (BB) P2 (bb) F1 (Bb) Same Band PositionBb Indistinguishable from BB Heterozygote Masked by Dominance

Figure: Co-Dominant vs. Dominant Gel Patterns. A co-dominant marker produces three distinguishable band profiles across homozygotes P1, P2, and heterozygote F1. A dominant marker only reports presence or absence of a single band, so the heterozygote F1 (Bb) is visually identical to homozygote P1 (BB).

2.3 Hybridization-Based Markers: RFLPs

Restriction Fragment Length Polymorphism (RFLP) was the first widely used hybridization-based DNA marker system (developed in 1975). RFLPs arise when point mutations, insertions, deletions, or inversions create or destroy restriction endonuclease recognition sites in genomic DNA, altering the length of restriction fragments produced after cleavage.

Molecular Mechanism of RFLP FormationAllele 1 (Wild-Type — Both EcoRI Sites Present) EcoRI EcoRI Fragment 1A — 2.0 kb Fragment 1B — 3.0 kbAllele 2 (Mutant — Site 2 Destroyed by Point Mutation) EcoRI Site Lost Fragment 2 — 5.0 kb (Single Large Fragment)

Figure: Molecular Basis of RFLP Formation. Allele 1 retains both EcoRI recognition sites and yields two fragments (2.0 kb and 3.0 kb). A point mutation in Allele 2 destroys the second EcoRI site, so cleavage produces a single, larger 5.0 kb fragment instead.

RFLP Characteristics

  • Inheritance Mode

    Co-dominant (both parental alleles are visible as distinct bands in F1 heterozygotes).

  • Maximum Heterozygosity

    0.5 (due to its diallelic nature: site present vs. site absent).

  • Requirements

    Requires high-quality, non-degraded genomic DNA (2–10 µg) and radiolabeled locus-specific cDNA probes.

Clinical Application

Prenatal Diagnosis of Sickle-Cell Anemia. Sickle-cell anemia is caused by a single nucleotide substitution (A → T) in the 6th codon of the β-globin gene (GAG → GTG), altering Glutamic Acid to Valine (E6V). This exact point mutation destroys a recognition site for the restriction enzymes MstII and CvnI (5′-CCTNAGG-3′).

MstII Restriction Map of the β-Globin GeneNormal βA Allele Site 1 Site 2 Site 3 0.2 kb 1.1 kb (Produces 0.2 kb and 1.1 kb fragments)Mutant βS Allele (Site 2 Destroyed) Site 1 Site Lost Site 3 1.3 kb (Single Fragment)

Figure: MstII Restriction Map of the β-Globin Gene. The normal βA allele retains all three MstII sites and yields 0.2 kb and 1.1 kb fragments. The sickle-cell βS mutation destroys Site 2, fusing the region into a single 1.3 kb fragment.

Pedigree & RFLP Southern Blot Analysis

In a prenatal screening for sickle-cell anemia, genomic DNA extracted from family members and fetal amniocytes is digested with MstII, separated by gel electrophoresis, and hybridized with a β-globin exon probe:

Family Pedigree & MstII Southern Blot Analysis Unaffected Carrier Affected I-1 (Father) βA/βS — Carrier I-2 (Mother) βA/βS — Carrier II-1 (Child 1) βA/βA — Normal II-2 (Child 2) βS/βS — Affected II-3 (Fetus) βA/βS — Carrier MstII Southern Blot — Band PatternsI-1 (Father) I-2 (Mother) II-1 (Child 1) II-2 (Child 2) II-3 (Fetus) 1.3 kb 1.1 kb βA/βS (Carrier) βA/βS (Carrier) βA/βA (Normal) βS/βS (Affected) βA/βS (Carrier) 1.3 kb — Mutant βS allele 1.1 kb — Normal βA allele

Figure: Pedigree and Southern Blot Analysis. Both parents (I-1, I-2) are heterozygous carriers showing both bands. Child 1 (II-1) is homozygous normal (1.1 kb only); Child 2 (II-2) is affected (1.3 kb only). The fetus (II-3) shows both bands, confirming an unaffected heterozygous carrier genotype.

  • Parents (I-1, I-2)

    Heterozygous carriers (βAβS) display both the 1.1 kb normal band and the 1.3 kb mutant band.

  • Child 1 (II-1)

    Homozygous normal (βAβA) displays only the 1.1 kb band.

  • Child 2 (II-2)

    Affected homozygous (βSβS) displays only the 1.3 kb band.

  • Fetus (II-3)

    Displays two bands (1.3 kb and 1.1 kb), confirming the fetus is an unaffected heterozygous carrier (βAβS).

2.4 Arbitrary Primed PCR Markers: RAPD & AFLP

2.4.1 RAPD (Random Amplification of Polymorphic DNA)

Principle: Uses a single short arbitrary oligonucleotide primer (8–12 nucleotides in length, 60% GC content) to amplify random genomic regions without requiring prior genome sequence knowledge.

Mechanism: Amplification occurs only when two primer-binding sites lie on opposite strands in an inverted orientation within a PCR-amplifiable distance (<3000 bp).

Molecular Mechanism of RAPD AmplificationVariety A — Amplification Occurs (Both Primer Sites Present) Primer (1) Primer (2) PCR Product AVariety B — No Amplification (Primer Site 2 Mutated / Lost) Primer (1) Site Lost NO PCR Product for Variety B

Figure: RAPD Amplification Mechanism. In Variety A, two inverted primer-binding sites lie within an amplifiable distance, producing a defined PCR product. In Variety B, a mutation eliminates the second primer site, so no product is amplified — this presence/absence difference is the RAPD polymorphism.

Limitations: Dominant marker (cannot distinguish BB from Bb); highly sensitive to reaction conditions (Mg2+ concentration, annealing temperature, template purity), leading to poor inter-laboratory reproducibility.

2.4.2 AFLP (Amplified Fragment Length Polymorphism)

AFLP combines the high reproducibility of restriction digestion (RFLP) with the high throughput of PCR. It does not require prior genomic sequence knowledge.

AFLP Multi-Step Procedure Genomic DNA 1. Dual Restriction Digestion (EcoRI + MseI) Overlapping Restriction Fragments 2. Ligation of EcoRI and MseI Adaptors Adaptor-Ligated Fragments (Sites Modified — Re-Cleavage Prevented) 3. Preamplification (+1 Selective Nucleotide) Subset of Fragments (Complexity Reduced by 1/16) 4. Selective PCR Amplification (+3 Selective Nucleotides) Highly Specific Polymorphic Band Pattern (50–100 Bands Resolved on Polyacrylamide Gel) Preamplification: 4¹ = 4× reduction per side (16× total) Selective PCR: 4³ × 4³ = 64 × 64 = 4,096× reduction Inheritance: Dominant marker — extremely high reliability and reproducibility

Figure: AFLP Procedure Overview. Genomic DNA is dually digested, ligated to adaptors, preamplified with one selective base, and finally amplified with three selective bases — progressively narrowing thousands of fragments down to 50–100 clean, highly polymorphic bands.

1Dual Digestion

Genomic DNA is cleaved simultaneously with a rare cutter (6-bp recognition enzyme, e.g., EcoRI, 5′-G^AATTC-3′) and a frequent cutter (4-bp recognition enzyme, e.g., MseI, 5′-T^TAA-3′).

2Adaptor Ligation

Synthetic double-stranded oligonucleotide adaptors complementary to the EcoRI and MseI overhangs are ligated to the fragment ends. Adaptor design alters the original recognition site so that restriction re-cleavage does not occur.

3Preamplification

PCR is performed using primers complementary to the adaptor sequences plus one extra selective 3′ base (e.g., +1 base, extending primer by A). This reduces template complexity by 41 = 4× per side (16× total).

4Selective Amplification

A second PCR is performed using stringent annealing conditions with primers bearing three selective 3′ base extensions (e.g., +3 bases). This reduces fragment complexity by 43 × 43 = 64 × 64 = 4,096×, yielding 50–100 clean, distinct bands resolved on polyacrylamide gels.

5Inheritance

Dominant marker with extremely high reliability and reproducibility.

2.5 Sequence-Specific SSLPs and SNPs

2.5.1 SSLPs (Simple Sequence Length Polymorphisms)

SSLPs are tandemly repeated DNA sequence arrays that vary in the total number of repeat units between individuals.

1Minisatellites (VNTRs — Variable Number of Tandem Repeats)
  • Repeat unit length: 10–100 base pairs.
  • Genomic distribution: Clustered predominantly in sub-telomeric regions.
  • Detection: Southern hybridization or long-range PCR.
2Microsatellites (SSRs / STRs — Simple Sequence Repeats)
  • Repeat unit length: 2–6 base pairs (e.g., (CA)n or (GATA)n).
  • Genomic distribution: Abundant and uniformly scattered throughout the nuclear genome.
  • Detection: PCR using specific primers flanking the hypervariable repeat region, followed by high-resolution PAGE separation.
  • Inheritance: Co-dominant, highly polymorphic, highly reproducible.
Microsatellite (SSR/STR) PCR AmplificationAllele 1 — 5 Repeat Units Forward Primer (CA)(CA)(CA)(CA)(CA) Reverse Primer Short PCR ProductAllele 2 — 8 Repeat Units Forward Primer (CA)(CA)(CA)(CA)(CA)(CA)(CA)(CA) Reverse Primer Long PCR ProductFragment length difference directly reflects repeat-number polymorphism between alleles

Figure: Microsatellite PCR Amplification. Primers flanking a (CA)n repeat amplify alleles of different lengths depending on repeat number — 5 repeats yield a short product, 8 repeats yield a proportionally longer product, resolved by high-resolution PAGE.

2.5.2 SNPs (Single Nucleotide Polymorphisms)

  • Definition

    Single base pair variations (A/G, C/T) at a specific genomic position present in at least 1% of a population.

  • Abundance

    The most abundant molecular marker in plant and animal genomes (occurring every 100–300 bp in human DNA).

  • Detection

    High-throughput DNA sequencing, TaqMan assays, matrix-assisted laser desorption/ionization mass spectrometry (MALDI-TOF), or allele-specific PCR.

  • Inheritance

    Co-dominant, diallelic.

2.5.3 Summary Comparison Matrix of DNA Marker Systems

FeatureRFLPRAPDAFLPSSR (Microsatellite)SNP
Detection BasisSouthern HybridizationPCR AmplificationPCR AmplificationPCR AmplificationHigh-throughput Assay / Sequencing
Prior Sequence DataRequired (cDNA probes)Not RequiredNot RequiredRequired (Flanking primers)Required (Locus sequence)
Inheritance ModeCo-dominantDominantDominantCo-dominantCo-dominant
Polymorphism LevelMediumMedium-HighVery HighExtremely HighHigh (Genome-wide)
DNA Quality/QuantityHigh Quality (5–10 µg)Low Quality (10–25 ng)Medium (100–500 ng)Low Quality (10–50 ng)Ultra-low (1–5 ng)
ReproducibilityHighLow-IntermediateHighHighExtremely High
Genome Mapping & DNA Forensics

Genome Mapping & DNA Forensics

Restriction Mapping Logic · Radiation Hybrid Mapping · DNA Profiling & Paternity Analysis

3. Genome Mapping Techniques

Genome mapping identifies the relative order, spacing, and physical coordinates of genes and molecular markers along chromosomes. Maps are divided into three distinct categories.

Genome Map Classifications Genome Maps Genetic Maps (Recombination Frequency measured in Centimorgans) 1% recombination = 1 cM = 1 Map Unit (MU) Cytological Maps (Chromosome Banding & Microscopic Landmarks) e.g., Polytene bands, Giemsa G-banding Physical Maps (Direct Physical Distance measured in Base Pairs) Restriction, STS & Radiation Hybrid mapping

Figure: Genome Map Classifications. Genetic maps use recombination frequency (centimorgans), cytological maps use microscopic chromosome banding, and physical maps measure absolute nucleotide distance in base pairs.

  • Map Type
    Genetic Maps

    Show marker order and relative linkage distances calculated from meiotic recombination frequencies (1% recombination = 1 centimorgan, cM = 1 Map Unit, MU).

  • Map Type
    Cytological Maps

    Map markers relative to microscopic banding patterns on stained chromosomes (e.g., polytene chromosome bands in Drosophila or Giemsa G-banding).

  • Map Type
    Physical Maps

    Measure absolute nucleotide distances between markers directly in base pairs (bp, kb, Mb). Common approaches include restriction mapping, Sequence-Tagged Site (STS) mapping, and Radiation Hybrid (RH) mapping.

3.1 Linear Restriction Mapping Logic

Restriction mapping determines the relative locations and distances between cleavage sites of different restriction enzymes along a DNA molecule.

Solved Analytical Problem 3

Double-Digest Linear Restriction Mapping. A 4.0-kb linear double-stranded DNA fragment is subjected to single and double digests using the restriction enzymes HindIII and BamHI. Digestion products are separated by agarose gel electrophoresis, yielding the following fragment sizes:

  • HindIII Digest: 2.8 kb, 1.2 kb
  • BamHI Digest: 1.8 kb, 1.3 kb, 0.9 kb
  • HindIII + BamHI Double Digest: 1.8 kb, 1.0 kb, 0.9 kb, 0.3 kb

Task: Determine the unique order and spatial coordinates of the HindIII and BamHI restriction cleavage sites along this 4.0-kb DNA molecule.

Mock Agarose Gel Electrophoresis PatternHindIII Digest HindIII + BamHI BamHI Digest 2.8 kb 1.8 kb 1.3 kb 1.2 kb 1.0 kb 0.9 kb 0.3 kb 1.8 kb & 0.9 kb bands are unchanged between BamHI and the double digest — HindIII does not cut inside them

Figure: Mock Gel Pattern for Problem 3. The HindIII digest produces 2 bands, the BamHI digest produces 3 bands, and the double digest produces 4 bands. The 1.8 kb and 0.9 kb BamHI bands persist unchanged in the double digest, while the 1.3 kb BamHI band splits into 1.0 kb + 0.3 kb.

1Analyze Single Digests
  • Total length of DNA = 2.8 + 1.2 = 4.0 kb.
  • HindIII cuts once (producing 2 fragments: 2.8 kb and 1.2 kb).
  • BamHI cuts twice (producing 3 fragments: 1.8 + 1.3 + 0.9 = 4.0 kb).
2Compare Single vs. Double Digest Fragments
  • The double digest yields 4 fragments totaling 1.8 + 1.0 + 0.9 + 0.3 = 4.0 kb.
  • The 1.8 kb and 0.9 kb fragments appear unchanged in both the single BamHI digest and the double digest — this proves HindIII does not cut within either of them.
  • Therefore, the HindIII cleavage site must lie inside the remaining 1.3 kb BamHI fragment.
  • Cleaving the 1.3 kb BamHI fragment with HindIII splits it into 1.0 kb + 0.3 kb = 1.3 kb.
3Determine Boundary Overlaps & Orientation
  • HindIII cuts the intact molecule into 2.8 kb and 1.2 kb.
  • To form the 2.8 kb HindIII fragment, combine the 1.8 kb BamHI fragment with the 1.0 kb sub-fragment (1.8 + 1.0 = 2.8 kb).
  • To form the 1.2 kb HindIII fragment, combine the 0.9 kb BamHI fragment with the 0.3 kb sub-fragment (0.9 + 0.3 = 1.2 kb).
Final Linear Restriction Map 0.0 kb Left End 1.8 kb BamHI Site 1 2.8 kb HindIII Site 3.1 kb BamHI Site 2 4.0 kb Right End 1.8 kb Fragment (BamHI) 1.0 kb Sub-block 0.3 kb Sub-block 0.9 kb Fragment (BamHI)

Figure: Final Linear Restriction Map. The unique site order along the 4.0 kb molecule is: Left End (0.0 kb) → BamHI Site 1 (1.8 kb) → HindIII Site (2.8 kb) → BamHI Site 2 (3.1 kb) → Right End (4.0 kb).

Restriction Sites: 0.0 kb (Left End) · 1.8 kb (BamHI Site 1) · 2.8 kb (HindIII Site) · 3.1 kb (BamHI Site 2) · 4.0 kb (Right End)

3.2 Circular Restriction Mapping (Plasmids)

Solved Analytical Problem 4

Circular Plasmid Mapping. A small circular plasmid DNA is cleaved with several restriction enzymes. Gel electrophoresis reveals the following fragment sizes:

  • EcoRI: 1.3 kb, 1.3 kb
  • HpaII: 2.6 kb
  • HindIII: 2.6 kb
  • EcoRI + HpaII: 1.3 kb, 0.8 kb, 0.5 kb
  • EcoRI + HindIII: 1.3 kb, 0.7 kb, 0.6 kb

Task: Is the original DNA molecule linear or circular? Prove mathematically, then construct a complete restriction map showing distances between cleavage sites.

1Proof of Circular Geometry
  • Summed fragments: EcoRI = 1.3 + 1.3 = 2.6 kb; HpaII = 2.6 kb (single cut); HindIII = 2.6 kb (single cut); EcoRI + HpaII = 1.3 + 0.8 + 0.5 = 2.6 kb.
  • In a linear molecule, n restriction cuts yield n+1 fragments. In a circular molecule, n cuts yield exactly n fragments.
  • EcoRI yields 2 fragments (n=2); HpaII and HindIII each yield only 1 fragment equal to the full length (n=1).
  • Therefore, the molecule is CIRCULAR with a total length of 2.6 kb.
2Locating EcoRI Sites

The two EcoRI sites lie directly opposite each other, dividing the 2.6 kb circle into two equal 1.3 kb arcs (coordinates: 0.0 kb / 2.6 kb and 1.3 kb).

3Locating the HpaII Site

Double digest EcoRI + HpaII yields 1.3 kb, 0.8 kb, and 0.5 kb. One 1.3 kb EcoRI arc remains completely intact, proving HpaII cuts inside the other 1.3 kb arc, dividing it into 0.8 kb and 0.5 kb. Position HpaII at coordinate 0.5 kb (or 0.8 kb clockwise from Site 1).

4Locating the HindIII Site

Double digest EcoRI + HindIII yields 1.3 kb, 0.7 kb, and 0.6 kb. Again, one 1.3 kb arc remains intact, while the other is split into 0.7 kb and 0.6 kb. Position HindIII at coordinate 0.6 kb relative to Site 1.

Circular Plasmid Restriction Map (2.6 kb Total) EcoRI Site 1 (0.0 / 2.6 kb) EcoRI Site 2 (1.3 kb) HpaII (0.5 kb) HindIII (0.6 kb)2.6 kb Total LengthCoordinates measured clockwise from EcoRI Site 1 (0.0 kb)

Figure: Circular Plasmid Restriction Map. Two EcoRI sites split the 2.6 kb circle into two equal 1.3 kb arcs. Within one arc, HpaII cuts at 0.5 kb and HindIII cuts at 0.6 kb from EcoRI Site 1, both determined independently via separate double digests.

3.3 Advanced Physical Mapping: STS and Radiation Hybrids

3.3.1 STS (Sequence-Tagged Site) Mapping

An STS is a short (200–500 bp) unique single-copy genomic sequence whose exact base composition and chromosomal location are known. STSs are detected specifically via PCR. An EST (Expressed Sequence Tag) is a specialized STS derived from a cDNA clone representing an actively expressed gene. STS mapping orders chromosome fragments by testing for the co-presence of shared STS markers.

3.3.2 Radiation Hybrid (RH) Mapping

Radiation hybrid mapping is a high-resolution physical mapping technique used to order STS markers along human chromosomes without requiring genetic polymorphism.

Radiation Hybrid Mapping Flow Human Somatic Cell (Target Chromosome) Lethal X-Ray / Gamma Irradiation (3,000–8,000 rads) Fragmented Chromosomes (Random Fragments, 5–10 Mb) Fusion with Non-Irradiated Hamster Host Cells (PEG) Rodent–Human Hybrid Cells (Contain Randomly Retained Human Fragments) Selection on HAT Media & STS PCR Screening (80–100 Clones) Co-Retention Probability Calculation (Closer STSs = Higher Co-Retention Frequency, Measured in Centirays, cR) Principle: Higher irradiation-breakage probability between distant markers lowers co-retention

Figure: Radiation Hybrid Mapping Workflow. Lethal irradiation fragments human chromosomes, which are rescued by fusion with hamster cells. Surviving hybrid clones retain random human fragments; screening a large clone panel by STS-PCR and measuring co-retention frequency orders markers by physical proximity.

  • Irradiation

    Human donor cells containing the chromosome of interest are exposed to lethal doses of X-rays or gamma radiation (3,000–8,000 rads), shattering the human chromosomes into random fragments (5–10 Mb).

  • Cell Fusion

    The irradiated dying human cells are fused with non-irradiated hamster (rodent) mutant cells lacking hypoxanthine-guanine phosphoribosyltransferase (HGPRT) using polyethylene glycol (PEG).

  • Selection & Panel Screening

    Hybrid cells are grown on HAT selective media. Surviving rodent-human hybrid clones retain a random subset of human chromosome fragments integrated into rodent chromosomes.

  • Co-Retention Analysis

    A panel of 80–100 independent radiation hybrid clones is screened by PCR for the presence or absence of specific STS markers.

  • Principle

    The closer two STS markers lie on a human chromosome, the lower the probability that X-ray irradiation will break the DNA between them, and the higher their co-retention frequency in the same hybrid clone. Retention distance is measured in Centirays (cR).

4. DNA Profiling and Forensics

DNA profiling (DNA typing or DNA fingerprinting), developed in 1984 by Sir Alec Jeffreys, exploits hypervariable non-coding repetitive loci (VNTRs and STRs) to establish positive identification and biological relationships between individuals.

Forensic DNA Profiling Mechanics (VNTR/STR Locus) Parent 1 A1 = 6 Repeats A2 = 9 Repeats Parent 2 B1 = 5 Repeats B2 = 7 RepeatsInheritance Pattern in Progeny (Mendelian Co-Dominance) Child 1: A2 (9) + B2 (7) Pattern: [9, 7] Child 2: A1 (6) + B2 (7) Pattern: [6, 7] Child 3: A1 (6) + B1 (5) Pattern: [6, 5]

Figure: Forensic DNA Profiling Mechanics. Each parent is heterozygous at a VNTR/STR locus. Each child inherits exactly one allele from each parent, producing a distinct co-dominant band pattern that reflects the Mendelian combination transmitted.

4.1 Solved Paternity & Forensic Pedigree Problem

Case File

A disputed paternity case involves a Mother, a Child, and two Alleged Fathers (AF1 and AF2). Genomic DNA is extracted, digested with a restriction enzyme, and hybridized with a single-locus VNTR probe. The resulting Southern blot band sizes (in kilobases, kb) are analyzed below.

Paternity Southern Blot DataMother Child Alleged Father 1 Alleged Father 210.5 kb 8.2 kb 6.4 kb 5.1 kb 3.8 kb Maternal allele Obligate paternal allele (10.5 kb) Non-matching (excludes AF1)AF2 shares the obligate 10.5 kb paternal band with the Child — AF1 does not

Figure: Paternity Southern Blot. The Mother contributes the maternal allele to the Child. The Child's remaining band (10.5 kb) is the obligate paternal allele — present in AF2 but absent in AF1.

1Identify the Maternal Alleles in the Child

Mother's genotype bands: 8.2 kb and 5.1 kb. Child's genotype bands: 10.5 kb and 6.4 kb. In a standard single-locus co-dominant VNTR, the Child must inherit exactly one allele from the Mother and one allele from the biological Father.

2Re-Evaluate the Band Alignment
  • Child has 10.5 kb and 6.4 kb.
  • Alleged Father 1 (AF1) has bands at 6.4 kb and 3.8 kb.
  • Alleged Father 2 (AF2) has bands at 10.5 kb and 5.1 kb.
3Determine the Obligate Paternal Allele

Given the Mother's allele profile, the non-maternal obligate paternal allele in the Child is 10.5 kb. Candidate AF1 possesses the 6.4 kb band but lacks the obligate 10.5 kb paternal band. Candidate AF2 possesses the obligate 10.5 kb paternal band transmitted to the Child.

Conclusion: Alleged Father 2 (AF2) is the biological father. Alleged Father 1 is excluded because he cannot supply the obligate 10.5 kb paternal band present in the Child's DNA profile.

5. Summary Matrix of Genomic Analysis Techniques

TechniqueFull NamePrimary Purpose / ApplicationKey AdvantageMajor Limitation
Genomic LibraryWhole-Genome Clone LibraryCloning total genomic sequence (coding & non-coding)Complete representation of entire genomeHigh screening effort due to vast intron/repeat background
cDNA LibraryComplementary DNA LibraryExpression of functional eukaryotic proteins in host cellsIntron-free, coding sequences onlyTissue-specific, lacks non-expressed regulatory genes
RFLPRestriction Fragment Length PolymorphismCo-dominant molecular mapping & disease diagnosisHigh reproducibility, co-dominantRequires large amounts (>5 µg) of high-quality DNA
RAPDRandom Amplification of Polymorphic DNAFast genetic diversity screeningNo prior sequence data requiredLow inter-laboratory reproducibility, dominant marker
AFLPAmplified Fragment Length PolymorphismHigh-density genomic fingerprintingExtremely high band output and reliabilityDominant marker, technically demanding multi-step protocol
SSR / STRSimple Sequence Repeat (Microsatellite)Pedigree analysis, MAS, linkage mappingHighly polymorphic, co-dominant, PCR-basedRequires locus-specific primer design
SNPSingle Nucleotide PolymorphismHigh-throughput genome-wide association studies (GWAS)Most abundant marker in genome, scalableDiallelic nature provides lower information per locus than SSRs
Radiation HybridRadiation Hybrid MappingHigh-resolution physical mapping of STSsBypasses need for genetic polymorphisms or meiotic crossesRequires specialized hamster-human cell lines and HAT selection

In this lesson

Scroll to Top