Close

Protein Sequencing: Techniques & Applications

Introduction Edman Degradation Sanger's Method Mass Spectrometry Dansyl Chloride Method De Novo Sequencing Next Generation High Throughput Applications

Introduction to Protein Sequencing

What is Protein Sequencing?

Protein sequencing is the process of determining the amino acid sequence of a protein or peptide. Proteins are polymers of amino acids linked by peptide bonds, and the precise order of these amino acids dictates the protein's three-dimensional structure, function, and interactions within a biological system. Understanding this sequence is paramount for comprehending protein function, disease mechanisms, and for the development of novel therapeutics and biotechnological products.

History of Protein Sequencing

The journey of protein sequencing began in the mid-20th century, marking a significant leap in our understanding of biological molecules. The pioneering work laid the foundation for the sophisticated techniques available today.

Edman Degradation Method for Protein Sequencing

The Edman degradation method, developed by Pehr Edman, is a classical technique for sequential N-terminal protein sequencing. It involves a series of chemical reactions that cleave one amino acid at a time from the N-terminus of a peptide or protein.

The process is cyclic and typically involves three main steps:

Overview of single molecule protein fluorosequencing. (OA Literature) Fig 1. Overview of single molecule protein fluorosequencing with various potential sources of errors highlighted in red.1

After identification of the N-terminal amino acid, the cycle can be repeated on the shortened peptide, allowing for the sequential determination of the amino acid sequence.

Table 1. Advantages and Limitations of Edman Degradation

Feature Advantage Limitation
Throughput Automated sequencers allow for relatively fast sequencing of multiple residues. Limited to ~50-60 amino acids due to accumulating impurities and decreasing yields.
Accuracy High accuracy for N-terminal sequence determination. Not suitable for highly modified proteins or blocked N-termini.
Cost Relatively inexpensive per residue once the instrument is acquired. Requires specialized reagents and equipment.
Sample Size Requires microgram quantities of protein. Can be affected by contaminants.

Sanger's Method for Protein Sequencing

While primarily known for his groundbreaking work on insulin sequencing, Frederick Sanger also developed methods for identifying N-terminal amino acids. His approach to N-terminal determination involved the use of 2,4-dinitrofluorobenzene (DNFB), also known as Sanger's reagent.

The Sanger method for N-terminal amino acid identification operates on the principle that DNFB reacts with the free α-amino group of the N-terminal amino acid under mild alkaline conditions, forming a dinitrophenyl (DNP) derivative. After acid hydrolysis of the protein, the N-terminal DNP-amino acid is released and can be identified chromatographically. However, this method only identifies the N-terminal amino acid and destroys the rest of the peptide chain, making it unsuitable for sequential degradation like Edman degradation. Therefore, its application in full protein sequencing is limited to identifying the very first amino acid.

Mass Spectrometry-Based Methods for Protein Sequencing

Mass spectrometry (MS) has revolutionized protein sequencing due to its high sensitivity, speed, and ability to characterize complex protein mixtures and post-translational modifications.

Protein Sequencing by Mass Spectrometry

The fundamental principle of MS-based protein sequencing involves:

For protein sequencing, proteins are typically digested into smaller peptides using proteolytic enzymes like trypsin. These peptides are then introduced into the mass spectrometer.

Tandem Mass Spectrometry Protein Sequencing

Tandem mass spectrometry (MS/MS) is the workhorse of modern protein sequencing. It involves two or more stages of mass analysis, allowing for fragmentation of selected ions and subsequent analysis of the fragment ions. This fragmentation pattern provides the sequence information.

The general workflow for peptide sequencing by MS/MS is as follows:

Protein Sequencing and Identification Using Tandem Mass Spectrometry

MS/MS data can be used for both de novo sequencing (determining the sequence without prior knowledge) and database searching (identifying proteins by matching experimental spectra to theoretical spectra derived from known protein sequences).

Table 2: Common Fragmentation Methods in Tandem Mass Spectrometry

Method Principle Advantages Disadvantages
CID Ions collide with neutral gas, gaining internal energy. Robust, commonly available. Can lead to neutral losses (e.g., H2O, NH3), complex spectra, limited for PTMs.
HCD Fragmentation within the collision cell, higher energy. More predictable fragmentation, good for quantitative proteomics. Still primarily cleaves backbone bonds.
ETD Electron transfer to peptide ions, non-ergodic fragmentation. Preserves labile post-translational modifications (PTMs), good for large peptides. Requires specialized instrumentation, lower fragmentation efficiency for some peptides.

Dansyl Chloride Method for Protein Sequencing

Similar to Sanger's reagent, dansyl chloride (5-dimethylamino-1-naphthalenesulfonyl chloride) is another reagent used for N-terminal amino acid determination. It reacts with the free α-amino group of the N-terminal amino acid to form a fluorescent dansyl derivative.

The process involves:

While highly sensitive due to the fluorescence of the dansyl derivative, this method, like Sanger's, is also degradative, meaning the entire peptide is hydrolyzed to identify only the N-terminal amino acid. It is not suitable for sequential sequencing beyond the first residue. Its main advantage was its significantly higher sensitivity compared to the DNP method, making it valuable for very small sample sizes in earlier times.

De Novo Protein Sequencing

De novo protein sequencing is the process of determining the amino acid sequence of a peptide or protein directly from its mass spectrometry fragmentation data, without relying on a pre-existing sequence database. This is a crucial technique when working with:

The process of de novo sequencing involves:

While powerful, de novo sequencing can be computationally intensive and may require higher quality and more complete MS/MS data than database searching for reliable results.

Next Generation Protein Sequencing

The term "Next Generation Protein Sequencing" (NGPS) is emerging, mirroring the revolution in DNA sequencing. While not yet as mature or widespread as NGS for DNA, NGPS aims to achieve ultra-high throughput, lower cost, and direct sequencing of intact proteins or long peptides, bypassing the need for extensive enzymatic digestion and chromatographic separation.

Current NGPS approaches are still largely in research and development phases but involve exciting technologies such as:

NGPS holds immense promise for enabling truly comprehensive proteome analysis, rapid biomarker discovery, and personalized medicine by providing unprecedented insights into the protein world.

High Throughput Protein Sequencing

High throughput protein sequencing refers to the ability to sequence a large number of proteins or peptides rapidly and efficiently. This is primarily achieved through:

High throughput protein sequencing is critical for large-scale proteomics studies, enabling the comparison of proteomes under different conditions (e.g., disease vs. healthy, drug-treated vs. untreated) and the discovery of novel protein biomarkers and drug targets.

Application of Protein Sequencing

Biopharmaceutical Characterization

Biomarker Discovery and Validation

Drug Target Identification

Fundamental Research

Enzyme Engineering and Protein Design

Food Science and Agriculture

Forensics and Toxicology

From unraveling the fundamental building blocks of life to driving the development of innovative therapies, the impact of protein sequencing is profound and ever-expanding. At Creative Biolabs, we combine decades of experience with cutting-edge technologies and a dedicated team of experts to provide comprehensive de novo sequencing services. We offer de novo antibody sequencing and de novo protein sequencing services, powered by our propriety DASS (Database Assisted Shotgun Sequencing) technology to meet the diverse protein research needs of our clients, driving innovation and advancement in the field of biomedical science.

Learn more about Creative Biolabs' de novo antibody sequencing services:

Reference
  1. Smith, Matthew Beauregard, et al. "Estimating error rates for single molecule protein sequencing experiments." PLOS Computational Biology 20.7 (2024): e1012258. Distributed under Open Access license CC BY 4.0, without modification. https://doi.org/10.1371/journal.pcbi.1012258

All listed services and products are For Research Use Only. Do Not use in any diagnostic or therapeutic applications.

Online Inquiry
CONTACT US
USA:
Europe:
Germany:
Call us at:
USA:
UK:
Germany:
Fax:
Email:
Our customer service representatives are available 24 hours a day, 7 days a week. Contact Us
© 2026 Creative Biolabs. | Contact Us