Close

Sanger Protein Sequencing: A Foundational Pillar in Molecular Biology

Introduction Step-by-Step Breakdown First Protein Sequenced Advantages & Disadvantages

Introduction to Sanger's Protein Sequencing

What is Sanger's Protein Sequencing Method?

Sanger's protein sequencing method, developed by British biochemist Frederick Sanger and his colleagues in the late 1940s and early 1950s, is a chemical technique designed to determine the linear sequence of amino acids in a polypeptide chain. The core principle involves selectively labeling the N-terminal amino acid of a protein or peptide, followed by the complete hydrolysis of the polypeptide into its constituent amino acids. The labeled N-terminal amino acid is then identified, typically by chromatography. While this method identifies only one amino acid per cycle (the N-terminal one), through iterative application on overlapping peptide fragments, the entire sequence of a protein could be painstakingly reconstructed.

The Historical Significance of Sequencing the First Protein

Prior to Sanger's work, the precise chemical nature of proteins was a subject of considerable debate. While it was known that proteins were composed of amino acids, the idea that each protein had a unique, genetically determined, and invariant sequence of these amino acids was not universally accepted. Some theories proposed proteins might be complex colloidal aggregates with no fixed structure.

Sanger's successful sequencing of insulin, completed in 1955, was a landmark achievement for several reasons:

The Sanger Method: A Step-by-Step Breakdown

The Sanger method, while laborious by modern standards, was elegant in its chemical logic for identifying the N-terminal amino acid of a polypeptide.

Principle of the Method: N-terminal Amino Acid Identification

The fundamental principle relies on the chemical derivatization of the free α-amino group present at the N-terminus of a polypeptide chain. This derivatization creates a stable, identifiable tag on the first amino acid. Subsequent hydrolysis breaks all peptide bonds, releasing all amino acids, but the N-terminal one remains tagged and can be isolated and identified.

Key Reagent: 2,4-dinitrofluorobenzene (DNFB)

The key reagent employed by Sanger was 1-fluoro-2,4-dinitrobenzene, often referred to as 2,4-dinitrofluorobenzene (DNFB) or Sanger's reagent.

DNFB is highly reactive towards nucleophilic groups, particularly primary and secondary amines, under mild conditions. The fluorine atom is a good leaving group, facilitating nucleophilic aromatic substitution.

Chemical Process

The polypeptide is reacted with DNFB, typically under mildly alkaline conditions (e.g., sodium bicarbonate solution, pH ~8-9). At this pH, the N-terminal α-amino group is largely deprotonated and thus highly nucleophilic. It attacks the carbon atom to which the fluorine is attached in DNFB, displacing the fluoride ion and forming a stable 2,4-dinitrophenyl (DNP) derivative of the N-terminal amino acid. This DNP-peptide is typically yellow due to the dinitrophenyl group.

Other free amino groups in the polypeptide, such as the ε-amino group of lysine side chains or a free N-terminus of another chain if the protein is multimeric, will also react with DNFB. However, only the α-amino group of the N-terminal amino acid will yield a DNP-amino acid that was originally at the N-terminus of that specific chain.

After the labeling reaction, the DNP-polypeptide is subjected to complete acid hydrolysis. This is typically achieved by heating the DNP-polypeptide in a strong acid solution (e.g., 6 M HCl at 100-110°C for several hours, typically 12-24 hours). This treatment cleaves all the peptide bonds within the polypeptide chain, releasing the individual amino acids.

The DNP group forms a very stable covalent bond with the N-terminal amino acid's α-nitrogen, a bond that is resistant to acid hydrolysis. Therefore, after hydrolysis, the N-terminal amino acid is recovered as its DNP-derivative (DNP-amino acid), while all other amino acids are released in their free, unmodified form.

The mixture resulting from hydrolysis contains free amino acids and the yellow DNP-amino acid (or DNP-derivatives of certain side chains like lysine, if present). The DNP-amino acid is then separated from the free amino acids and identified.

The First Protein Sequenced: Insulin

The successful application of this method to determine the complete amino acid sequence of bovine insulin was a monumental task that took Sanger and his team nearly a decade (from roughly 1945 to 1955).

The Structure of Insulin: Two Polypeptide Chains

Bovine insulin is a relatively small protein, but it presented complexities. It consists of two distinct polypeptide chains:

These two chains are covalently linked by two inter-chain disulfide bonds. The A-chain also contains one intra-chain disulfide bond.

Sanger's Strategy for Sequencing Insulin

Sanger's approach to sequencing insulin was systematic and meticulous:

  1. Separation of Chains: The A and B chains were first separated. This was achieved by oxidizing the disulfide bonds with performic acid, converting cysteine residues to cysteic acid. This prevented the disulfide bonds from reforming and allowed for the isolation of the individual linear chains.
  2. N-terminal Analysis of Intact Chains: Using the DNFB method, Sanger identified glycine (Gly) as the N-terminal amino acid of the A-chain and phenylalanine (Phe) as the N-terminal of the B-chain.
  3. Fragmentation and Sequencing of Peptides: Since the DNFB method only identifies the N-terminal residue before destroying the rest of the peptide during hydrolysis, sequencing longer chains required a fragmentation strategy:
    • Partial Hydrolysis: The separated A and B chains were subjected to partial acid hydrolysis or enzymatic digestion (e.g., with trypsin, chymotrypsin, pepsin) to generate a complex mixture of smaller, overlapping peptide fragments.
    • Fractionation of Peptides: These fragments were then painstakingly separated using techniques like paper chromatography and electrophoresis.
    • Sequencing of Individual Fragments: Each purified fragment was then subjected to N-terminal analysis using the DNFB method. For very short peptides (di- or tri-peptides), the sequence could often be deduced from the N-terminal residue and the overall amino acid composition.
    • Iterative N-terminal Analysis: For slightly longer fragments, after identifying the N-terminal DNP-amino acid, the remaining DNP-peptide could sometimes be further partially hydrolyzed, and a new N-terminal (now internal to the original fragment) could be identified.
  4. Ordering Peptides using Overlaps: By analyzing the sequences of these overlapping fragments, Sanger could deduce their original order in the intact A and B chains, much like assembling a jigsaw puzzle. For example, if fragment 1 was Gly-Ile-Val and fragment 2 was Val-Glu-Gln, the overlap (Val) would indicate a sequence of Gly-Ile-Val-Glu-Gln.
  5. Determination of Disulfide Bond Positions: After sequencing the linear A and B chains, the positions of the disulfide bonds were determined by hydrolyzing intact insulin (without prior oxidation of disulfide bonds) into fragments containing intact cystine (two cysteine residues linked by a disulfide bond). These cystine-containing peptides were isolated, and the linked cysteine residues were identified, thus pinpointing the connections between the chains and within the A-chain.

The Impact of Sequencing Insulin on Understanding Protein Structure and Function

Advantages and Disadvantages of Sanger's Protein Sequencing

Strengths of the Method

The Major Drawback: Inefficiency for Long Polypeptide Chains

The most significant limitation of the Sanger method was its destructive nature. In each cycle of labeling and hydrolysis:

Only the N-terminal amino acid was identified.

The main issues encountered in the reading of DNA chromatograms of PCR products by the Sanger sequencing method. (OA Literature) Fig. 1 The main issues encountered in the reading of DNA chromatograms of PCR products based on the Sanger sequencing method.1

The Tedious and Time-Consuming Nature of the Process

The Sanger method was incredibly labor-intensive:

Table 1. Comparison of Sanger's Method with Modern Techniques

Feature Sanger Method (DNFB) Edman Degradation (PITC) Mass Spectrometry (MS/MS)
Principle N-terminal labeling, full hydrolysis Sequential N-terminal cleavage Peptide fragmentation, mass analysis
Efficiency/Speed Very slow, laborious Moderate, automatable (sequenators) Very fast, high-throughput
Sensitivity Low (milligram to microgram) Moderate (nanomole to picomole) High (picomole to femtomole)
Length per Run 1 amino acid (N-terminal) Up to ~50 amino acids (automated) Full peptide sequence (typically 5-30 AA per peptide, many peptides analyzed)
Sample Destruction Complete peptide hydrolysis N-terminal residue removed, rest intact Sample consumed in ionization/fragmentation
PTM Analysis Very difficult, indirect Difficult, some PTMs block Edman Excellent, can identify and locate PTMs
Mixture Analysis Very difficult, requires pure peptides Difficult, requires pure peptides Good, can analyze complex mixtures (LC-MS/MS)
Automation None Yes (automated sequenators) Yes (autosamplers, LC, data analysis)
Primary Use Today Historically significant, educational N-terminal confirmation, specialized Dominant for most protein sequencing tasks

Sanger's protein sequencing method not only provided the first glimpse into the precise architecture of proteins but also laid the conceptual and methodological groundwork for decades of subsequent advancements in both protein and nucleic acid sequencing, fundamentally shaping the landscape of modern molecular biology. At Creative Biolabs, we combine decades of experience with cutting-edge technologies and a dedicated team of experts to provide comprehensive de novo sequencing services. We offer de novo antibody sequencing and de novo protein sequencing services, powered by our propriety DASS (Database Assisted Shotgun Sequencing) technology to meet the diverse protein research needs of our clients, driving innovation and advancement in the field of biomedical science.

Learn more about Creative Biolabs' de novo antibody sequencing services:

Reference
  1. Al-Shuhaib, Mohammed Baqur S., and Hayder O. Hashim. "Mastering DNA chromatogram analysis in Sanger sequencing for reliable clinical analysis." Journal of Genetic Engineering and Biotechnology 21.1 (2023): 115.Distributed under Open Access license CC BY 4.0, without modification. https://doi.org/10.1186/s43141-023-00587-6

All listed services and products are For Research Use Only. Do Not use in any diagnostic or therapeutic applications.

Online Inquiry
CONTACT US
USA:
Europe:
Germany:
Call us at:
USA:
UK:
Germany:
Fax:
Email:
Our customer service representatives are available 24 hours a day, 7 days a week. Contact Us
© 2026 Creative Biolabs. | Contact Us