Sanger Protein Sequencing: A Foundational Pillar in Molecular Biology
Introduction Step-by-Step Breakdown First Protein Sequenced Advantages & Disadvantages
Introduction to Sanger's Protein Sequencing
What is Sanger's Protein Sequencing Method?
Sanger's protein sequencing method, developed by British biochemist Frederick Sanger and his colleagues in the late 1940s and early 1950s, is a chemical technique designed to determine the linear sequence of amino acids in a polypeptide chain. The core principle involves selectively labeling the N-terminal amino acid of a protein or peptide, followed by the complete hydrolysis of the polypeptide into its constituent amino acids. The labeled N-terminal amino acid is then identified, typically by chromatography. While this method identifies only one amino acid per cycle (the N-terminal one), through iterative application on overlapping peptide fragments, the entire sequence of a protein could be painstakingly reconstructed.
The Historical Significance of Sequencing the First Protein
Prior to Sanger's work, the precise chemical nature of proteins was a subject of considerable debate. While it was known that proteins were composed of amino acids, the idea that each protein had a unique, genetically determined, and invariant sequence of these amino acids was not universally accepted. Some theories proposed proteins might be complex colloidal aggregates with no fixed structure.
Sanger's successful sequencing of insulin, completed in 1955, was a landmark achievement for several reasons:
-
It provided the first definitive proof that a protein has a precisely defined amino acid sequence.
-
It established that proteins are not random polymers but highly ordered molecules.
-
It laid the groundwork for understanding the relationship between a protein's primary structure and its three-dimensional conformation and, consequently, its biological function.
-
This monumental work earned Frederick Sanger his first Nobel Prize in Chemistry in 1958, recognizing "his work on the structure of proteins, especially that of insulin."
The Sanger Method: A Step-by-Step Breakdown
The Sanger method, while laborious by modern standards, was elegant in its chemical logic for identifying the N-terminal amino acid of a polypeptide.
Principle of the Method: N-terminal Amino Acid Identification
The fundamental principle relies on the chemical derivatization of the free α-amino group present at the N-terminus of a polypeptide chain. This derivatization creates a stable, identifiable tag on the first amino acid. Subsequent hydrolysis breaks all peptide bonds, releasing all amino acids, but the N-terminal one remains tagged and can be isolated and identified.
Key Reagent: 2,4-dinitrofluorobenzene (DNFB)
The key reagent employed by Sanger was 1-fluoro-2,4-dinitrobenzene, often referred to as 2,4-dinitrofluorobenzene (DNFB) or Sanger's reagent.
DNFB is highly reactive towards nucleophilic groups, particularly primary and secondary amines, under mild conditions. The fluorine atom is a good leaving group, facilitating nucleophilic aromatic substitution.
Chemical Process
-
Labeling the N-terminal Amino Acid
The polypeptide is reacted with DNFB, typically under mildly alkaline conditions (e.g., sodium bicarbonate solution, pH ~8-9). At this pH, the N-terminal α-amino group is largely deprotonated and thus highly nucleophilic. It attacks the carbon atom to which the fluorine is attached in DNFB, displacing the fluoride ion and forming a stable 2,4-dinitrophenyl (DNP) derivative of the N-terminal amino acid. This DNP-peptide is typically yellow due to the dinitrophenyl group.
Other free amino groups in the polypeptide, such as the ε-amino group of lysine side chains or a free N-terminus of another chain if the protein is multimeric, will also react with DNFB. However, only the α-amino group of the N-terminal amino acid will yield a DNP-amino acid that was originally at the N-terminus of that specific chain.
-
Hydrolysis of the Polypeptide Chain
After the labeling reaction, the DNP-polypeptide is subjected to complete acid hydrolysis. This is typically achieved by heating the DNP-polypeptide in a strong acid solution (e.g., 6 M HCl at 100-110°C for several hours, typically 12-24 hours). This treatment cleaves all the peptide bonds within the polypeptide chain, releasing the individual amino acids.
The DNP group forms a very stable covalent bond with the N-terminal amino acid's α-nitrogen, a bond that is resistant to acid hydrolysis. Therefore, after hydrolysis, the N-terminal amino acid is recovered as its DNP-derivative (DNP-amino acid), while all other amino acids are released in their free, unmodified form.
-
Identification of the Labeled Amino Acid via Chromatography
The mixture resulting from hydrolysis contains free amino acids and the yellow DNP-amino acid (or DNP-derivatives of certain side chains like lysine, if present). The DNP-amino acid is then separated from the free amino acids and identified.
-
Extraction: DNP-amino acids are soluble in organic solvents (e.g., ether or ethyl acetate), while free amino acids are more water-soluble. This allows for an initial separation by solvent extraction.
-
Chromatography: The extracted DNP-amino acids were typically identified using paper chromatography or, later, thin-layer chromatography (TLC). The DNP-amino acids are colored, making them visible on the chromatogram. Each DNP-amino acid has a characteristic migration rate (Rf value) in a given solvent system. By comparing the Rf value of the unknown DNP-amino acid with those of known DNP-amino acid standards run on the same chromatogram, the identity of the N-terminal amino acid could be determined.
The First Protein Sequenced: Insulin
The successful application of this method to determine the complete amino acid sequence of bovine insulin was a monumental task that took Sanger and his team nearly a decade (from roughly 1945 to 1955).
The Structure of Insulin: Two Polypeptide Chains
Bovine insulin is a relatively small protein, but it presented complexities. It consists of two distinct polypeptide chains:
-
A-chain: Comprising 21 amino acids.
-
B-chain: Comprising 30 amino acids.
These two chains are covalently linked by two inter-chain disulfide bonds. The A-chain also contains one intra-chain disulfide bond.
Sanger's Strategy for Sequencing Insulin
Sanger's approach to sequencing insulin was systematic and meticulous:
-
Separation of Chains: The A and B chains were first separated. This was achieved by oxidizing the disulfide bonds with performic acid, converting cysteine residues to cysteic acid. This prevented the disulfide bonds from reforming and allowed for the isolation of the individual linear chains.
-
N-terminal Analysis of Intact Chains: Using the DNFB method, Sanger identified glycine (Gly) as the N-terminal amino acid of the A-chain and phenylalanine (Phe) as the N-terminal of the B-chain.
-
Fragmentation and Sequencing of Peptides: Since the DNFB method only identifies the N-terminal residue before destroying the rest of the peptide during hydrolysis, sequencing longer chains required a fragmentation strategy:
-
Partial Hydrolysis: The separated A and B chains were subjected to partial acid hydrolysis or enzymatic digestion (e.g., with trypsin, chymotrypsin, pepsin) to generate a complex mixture of smaller, overlapping peptide fragments.
-
Fractionation of Peptides: These fragments were then painstakingly separated using techniques like paper chromatography and electrophoresis.
-
Sequencing of Individual Fragments: Each purified fragment was then subjected to N-terminal analysis using the DNFB method. For very short peptides (di- or tri-peptides), the sequence could often be deduced from the N-terminal residue and the overall amino acid composition.
-
Iterative N-terminal Analysis: For slightly longer fragments, after identifying the N-terminal DNP-amino acid, the remaining DNP-peptide could sometimes be further partially hydrolyzed, and a new N-terminal (now internal to the original fragment) could be identified.
-
Ordering Peptides using Overlaps: By analyzing the sequences of these overlapping fragments, Sanger could deduce their original order in the intact A and B chains, much like assembling a jigsaw puzzle. For example, if fragment 1 was Gly-Ile-Val and fragment 2 was Val-Glu-Gln, the overlap (Val) would indicate a sequence of Gly-Ile-Val-Glu-Gln.
-
Determination of Disulfide Bond Positions: After sequencing the linear A and B chains, the positions of the disulfide bonds were determined by hydrolyzing intact insulin (without prior oxidation of disulfide bonds) into fragments containing intact cystine (two cysteine residues linked by a disulfide bond). These cystine-containing peptides were isolated, and the linked cysteine residues were identified, thus pinpointing the connections between the chains and within the A-chain.
The Impact of Sequencing Insulin on Understanding Protein Structure and Function
-
Confirmation of Defined Primary Structure: It irrefutably demonstrated that proteins have a specific, covalent sequence of amino acids.
-
Genetic Basis of Protein Structure: It strongly supported the idea that the sequence of amino acids in a protein is genetically determined by the sequence of nucleotides in DNA (though the genetic code itself was yet to be deciphered).
-
Structure-Function Paradigm: It provided the foundation for the central paradigm that a protein's primary structure dictates its higher-order structures (secondary, tertiary, quaternary), which in turn determine its biological function.
-
Species Variation: Comparison of insulin sequences from different species later revealed conserved and variable regions, offering insights into evolution and the structural requirements for insulin's function.
-
Foundation for Peptide Synthesis: Knowing the exact sequence paved the way for the chemical synthesis of peptides and proteins.
Advantages and Disadvantages of Sanger's Protein Sequencing
Strengths of the Method
-
Unambiguous N-terminal Identification: It provided a clear and definitive way to identify the N-terminal amino acid of a polypeptide.
-
Pioneering Achievement: It was the first method that successfully sequenced a protein, opening up a new era in biochemistry.
-
Relatively Simple Reagents and Chemistry (for its time): While laborious, the underlying chemical reactions were understandable and manageable with the techniques available in the mid-20th century.
-
Robust Labeling: The DNP group formed a stable, colored derivative, which aided in its detection and isolation.
The Major Drawback: Inefficiency for Long Polypeptide Chains
The most significant limitation of the Sanger method was its destructive nature. In each cycle of labeling and hydrolysis:
Only the N-terminal amino acid was identified.
-
The rest of the polypeptide chain was completely hydrolyzed and thus "lost" for further sequential analysis from that same molecule.
-
This meant that to sequence an entire protein, it had to be broken down into many small, overlapping peptides, each of which had to be purified and then subjected to N-terminal analysis. Reconstructing the full sequence from these fragments was an intellectually demanding and time-consuming puzzle.
Fig. 1 The main issues encountered in the reading of DNA chromatograms of PCR products based on the Sanger sequencing method.1
The Tedious and Time-Consuming Nature of the Process
The Sanger method was incredibly labor-intensive:
-
Multiple Hydrolysis and Chromatography Steps: Each fragment required labeling, hydrolysis, extraction, and chromatographic separation.
-
Large Sample Quantities: The method required relatively large amounts of purified protein due to losses at each step and the sensitivity limits of detection.
-
Manual Operation: The entire process was manual, with no automation.
-
Data Interpretation: Assembling the final sequence from overlapping fragments was a complex analytical task. Sanger's work on insulin, a small protein, took approximately 10 years.
Table 1. Comparison of Sanger's Method with Modern Techniques
|
Feature
|
Sanger Method (DNFB)
|
Edman Degradation (PITC)
|
Mass Spectrometry (MS/MS)
|
|
Principle
|
N-terminal labeling, full hydrolysis
|
Sequential N-terminal cleavage
|
Peptide fragmentation, mass analysis
|
|
Efficiency/Speed
|
Very slow, laborious
|
Moderate, automatable (sequenators)
|
Very fast, high-throughput
|
|
Sensitivity
|
Low (milligram to microgram)
|
Moderate (nanomole to picomole)
|
High (picomole to femtomole)
|
|
Length per Run
|
1 amino acid (N-terminal)
|
Up to ~50 amino acids (automated)
|
Full peptide sequence (typically 5-30 AA per peptide, many peptides analyzed)
|
|
Sample Destruction
|
Complete peptide hydrolysis
|
N-terminal residue removed, rest intact
|
Sample consumed in ionization/fragmentation
|
|
PTM Analysis
|
Very difficult, indirect
|
Difficult, some PTMs block Edman
|
Excellent, can identify and locate PTMs
|
|
Mixture Analysis
|
Very difficult, requires pure peptides
|
Difficult, requires pure peptides
|
Good, can analyze complex mixtures (LC-MS/MS)
|
|
Automation
|
None
|
Yes (automated sequenators)
|
Yes (autosamplers, LC, data analysis)
|
|
Primary Use Today
|
Historically significant, educational
|
N-terminal confirmation, specialized
|
Dominant for most protein sequencing tasks
|
Sanger's protein sequencing method not only provided the first glimpse into the precise architecture of proteins but also laid the conceptual and methodological groundwork for decades of subsequent advancements in both protein and nucleic acid sequencing, fundamentally shaping the landscape of modern molecular biology. At Creative Biolabs, we combine decades of experience with cutting-edge technologies and a dedicated team of experts to provide comprehensive de novo sequencing services. We offer de novo antibody sequencing and de novo protein sequencing services, powered by our propriety DASS (Database Assisted Shotgun Sequencing) technology to meet the diverse protein research needs of our clients, driving innovation and advancement in the field of biomedical science.
Learn more about Creative Biolabs' de novo antibody sequencing services:
Reference
-
Al-Shuhaib, Mohammed Baqur S., and Hayder O. Hashim. "Mastering DNA chromatogram analysis in Sanger sequencing for reliable clinical analysis." Journal of Genetic Engineering and Biotechnology 21.1 (2023): 115.Distributed under Open Access license CC BY 4.0, without modification. https://doi.org/10.1186/s43141-023-00587-6
All listed services and products are For Research Use Only. Do Not use in any diagnostic or therapeutic applications.