Protein Sequencing: Techniques & Applications
Introduction Edman Degradation Sanger's Method Mass Spectrometry Dansyl Chloride Method De Novo Sequencing Next Generation High Throughput Applications
Introduction to Protein Sequencing
What is Protein Sequencing?
Protein sequencing is the process of determining the amino acid sequence of a protein or peptide. Proteins are polymers of amino acids linked by peptide bonds, and the precise order of these amino acids dictates the protein's three-dimensional structure, function, and interactions within a biological system. Understanding this sequence is paramount for comprehending protein function, disease mechanisms, and for the development of novel therapeutics and biotechnological products.
History of Protein Sequencing
The journey of protein sequencing began in the mid-20th century, marking a significant leap in our understanding of biological molecules. The pioneering work laid the foundation for the sophisticated techniques available today.
-
Early Milestones: The first protein to have its complete amino acid sequence determined was bovine insulin in 1953 by Frederick Sanger. This monumental achievement, for which he was awarded the Nobel Prize in Chemistry in 1958, demonstrated that proteins have a defined, rather than random, sequence.
-
Automated Techniques: Per Edman's development of the Edman degradation method in 1950, and its subsequent automation, revolutionized protein sequencing, making it more efficient and accessible.
-
Emergence of Mass Spectrometry: The late 20th and early 21st centuries witnessed the rapid rise of mass spectrometry (MS) as a powerful tool for protein analysis, ultimately transforming protein sequencing into a high-throughput discipline.
Edman Degradation Method for Protein Sequencing
The Edman degradation method, developed by Pehr Edman, is a classical technique for sequential N-terminal protein sequencing. It involves a series of chemical reactions that cleave one amino acid at a time from the N-terminus of a peptide or protein.
The process is cyclic and typically involves three main steps:
-
Coupling: The α-amino group of the N-terminal amino acid reacts with phenyl isothiocyanate (PITC) under mild alkaline conditions to form a phenylthiocarbamoyl (PTC) derivative.
-
Cleavage: The PTC-peptide is then treated with a strong anhydrous acid (e.g., trifluoroacetic acid) to cleave the N-terminal amino acid as an anilinothiazolinone (ATZ) derivative, leaving the remaining peptide chain intact.
-
Conversion: The unstable ATZ derivative is converted to a more stable phenylthiohydantoin (PTH) amino acid by treatment with aqueous acid. This PTH-amino acid can then be identified by various chromatographic techniques, most commonly high-performance liquid chromatography (HPLC).
Fig 1. Overview of single molecule protein fluorosequencing with various potential sources of errors highlighted in red.1
After identification of the N-terminal amino acid, the cycle can be repeated on the shortened peptide, allowing for the sequential determination of the amino acid sequence.
Table 1. Advantages and Limitations of Edman Degradation
|
Feature
|
Advantage
|
Limitation
|
|
Throughput
|
Automated sequencers allow for relatively fast sequencing of multiple residues.
|
Limited to ~50-60 amino acids due to accumulating impurities and decreasing yields.
|
|
Accuracy
|
High accuracy for N-terminal sequence determination.
|
Not suitable for highly modified proteins or blocked N-termini.
|
|
Cost
|
Relatively inexpensive per residue once the instrument is acquired.
|
Requires specialized reagents and equipment.
|
|
Sample Size
|
Requires microgram quantities of protein.
|
Can be affected by contaminants.
|
Sanger's Method for Protein Sequencing
While primarily known for his groundbreaking work on insulin sequencing, Frederick Sanger also developed methods for identifying N-terminal amino acids. His approach to N-terminal determination involved the use of 2,4-dinitrofluorobenzene (DNFB), also known as Sanger's reagent.
The Sanger method for N-terminal amino acid identification operates on the principle that DNFB reacts with the free α-amino group of the N-terminal amino acid under mild alkaline conditions, forming a dinitrophenyl (DNP) derivative. After acid hydrolysis of the protein, the N-terminal DNP-amino acid is released and can be identified chromatographically. However, this method only identifies the N-terminal amino acid and destroys the rest of the peptide chain, making it unsuitable for sequential degradation like Edman degradation. Therefore, its application in full protein sequencing is limited to identifying the very first amino acid.
Mass Spectrometry-Based Methods for Protein Sequencing
Mass spectrometry (MS) has revolutionized protein sequencing due to its high sensitivity, speed, and ability to characterize complex protein mixtures and post-translational modifications.
Protein Sequencing by Mass Spectrometry
The fundamental principle of MS-based protein sequencing involves:
-
Ionization: Converting molecules into gas-phase ions.
-
Mass Analysis: Separating ions based on their mass-to-charge ratio (m/z).
-
Detection: Measuring the abundance of each ion.
For protein sequencing, proteins are typically digested into smaller peptides using proteolytic enzymes like trypsin. These peptides are then introduced into the mass spectrometer.
Tandem Mass Spectrometry Protein Sequencing
Tandem mass spectrometry (MS/MS) is the workhorse of modern protein sequencing. It involves two or more stages of mass analysis, allowing for fragmentation of selected ions and subsequent analysis of the fragment ions. This fragmentation pattern provides the sequence information.
The general workflow for peptide sequencing by MS/MS is as follows:
-
Precursor Ion Selection (MS1): In the first mass analyzer, a specific peptide ion (precursor ion) from a mixture is selected based on its m/z.
-
Fragmentation: The selected precursor ion is then fragmented in a collision cell using inert gas (e.g., argon or nitrogen) through collision-induced dissociation (CID), higher-energy collisional dissociation (HCD), or electron-transfer dissociation (ETD). This breaks the peptide bonds.
-
Fragment Ion Analysis (MS2): The resulting fragment ions (product ions) are then analyzed in the second mass analyzer. The mass differences between consecutive fragment ions correspond to the masses of individual amino acid residues, thereby allowing for the reconstruction of the peptide sequence.
Protein Sequencing and Identification Using Tandem Mass Spectrometry
MS/MS data can be used for both de novo sequencing (determining the sequence without prior knowledge) and database searching (identifying proteins by matching experimental spectra to theoretical spectra derived from known protein sequences).
-
De Novo Sequencing: This approach relies solely on the fragment ion mass differences to deduce the amino acid sequence. It is particularly useful for novel proteins or peptides with unknown sequences, or for characterizing modified peptides.
-
Database Searching: This is the most common approach for protein identification. Experimental MS/MS spectra are searched against protein sequence databases using search algorithms. The algorithm calculates theoretical fragmentation patterns for all peptides in the database and compares them to the observed experimental spectra, assigning a score to indicate the likelihood of a match.
Table 2: Common Fragmentation Methods in Tandem Mass Spectrometry
|
Method
|
Principle
|
Advantages
|
Disadvantages
|
|
CID
|
Ions collide with neutral gas, gaining internal energy.
|
Robust, commonly available.
|
Can lead to neutral losses (e.g., H2O, NH3), complex spectra, limited for PTMs.
|
|
HCD
|
Fragmentation within the collision cell, higher energy.
|
More predictable fragmentation, good for quantitative proteomics.
|
Still primarily cleaves backbone bonds.
|
|
ETD
|
Electron transfer to peptide ions, non-ergodic fragmentation.
|
Preserves labile post-translational modifications (PTMs), good for large peptides.
|
Requires specialized instrumentation, lower fragmentation efficiency for some peptides.
|
Dansyl Chloride Method for Protein Sequencing
Similar to Sanger's reagent, dansyl chloride (5-dimethylamino-1-naphthalenesulfonyl chloride) is another reagent used for N-terminal amino acid determination. It reacts with the free α-amino group of the N-terminal amino acid to form a fluorescent dansyl derivative.
The process involves:
-
Reaction: Dansyl chloride reacts with the N-terminal amino acid of the peptide or protein under mild alkaline conditions to form a dansyl-peptide.
-
Hydrolysis: The dansyl-peptide is then hydrolyzed, typically with acid.
-
Identification: The released dansyl-amino acid is highly fluorescent and can be detected with high sensitivity using thin-layer chromatography (TLC) or HPLC.
While highly sensitive due to the fluorescence of the dansyl derivative, this method, like Sanger's, is also degradative, meaning the entire peptide is hydrolyzed to identify only the N-terminal amino acid. It is not suitable for sequential sequencing beyond the first residue. Its main advantage was its significantly higher sensitivity compared to the DNP method, making it valuable for very small sample sizes in earlier times.
De Novo Protein Sequencing
De novo protein sequencing is the process of determining the amino acid sequence of a peptide or protein directly from its mass spectrometry fragmentation data, without relying on a pre-existing sequence database. This is a crucial technique when working with:
-
Novel proteins: Proteins from unsequenced genomes or organisms.
-
Variants and Mutations: Identifying unexpected amino acid substitutions or deletions not present in databases.
-
Post-translational modifications (PTMs): Characterizing complex PTMs that alter the mass of amino acid residues.
-
Peptides with unknown sequences: For instance, in drug discovery or biomarker identification.
The process of de novo sequencing involves:
-
High-resolution MS/MS data acquisition: Acquiring accurate mass data for both precursor and fragment ions.
-
Spectral interpretation: Analyzing the mass differences between fragment ions (b-ions, y-ions, etc.) to deduce the sequence. This often requires specialized software that can interpret complex fragmentation patterns.
-
Manual validation: Expert manual validation is often required to confirm the accuracy of de novo sequences, especially for challenging spectra.
While powerful, de novo sequencing can be computationally intensive and may require higher quality and more complete MS/MS data than database searching for reliable results.
Next Generation Protein Sequencing
The term "Next Generation Protein Sequencing" (NGPS) is emerging, mirroring the revolution in DNA sequencing. While not yet as mature or widespread as NGS for DNA, NGPS aims to achieve ultra-high throughput, lower cost, and direct sequencing of intact proteins or long peptides, bypassing the need for extensive enzymatic digestion and chromatographic separation.
Current NGPS approaches are still largely in research and development phases but involve exciting technologies such as:
-
Nanopore Sequencing: Similar to DNA nanopore sequencing, this method involves passing proteins or peptides through a nanoscale pore. Changes in current or optical signals as different amino acids pass through the pore can potentially be used to identify them. Challenges include distinguishing all 20 amino acids based on their subtle electrical or optical properties, and ensuring controlled translocation of the protein.
-
Single-Molecule Real-Time Sequencing: This involves monitoring the incorporation of labeled amino acids into a growing polypeptide chain on a single molecule level, or directly "reading" the amino acid sequence of a protein by analyzing its interaction with a sensor.
-
Advanced Mass Spectrometry Platforms: While MS/MS is already a "next-gen" technology compared to Edman degradation, continuous advancements in MS hardware (e.g., faster scanning, higher resolution, improved fragmentation techniques) and software (e.g., more sophisticated de novo algorithms, machine learning for spectral interpretation) are pushing the boundaries of what's possible in protein sequencing.
NGPS holds immense promise for enabling truly comprehensive proteome analysis, rapid biomarker discovery, and personalized medicine by providing unprecedented insights into the protein world.
High Throughput Protein Sequencing
High throughput protein sequencing refers to the ability to sequence a large number of proteins or peptides rapidly and efficiently. This is primarily achieved through:
-
Automation: Automated liquid handling systems, robotic sample preparation, and automated mass spectrometers minimize manual intervention and increase sample processing capacity.
-
Advanced LC-MS/MS Systems: Coupling high-performance liquid chromatography (LC) with sophisticated mass spectrometers allows for the rapid separation and analysis of complex peptide mixtures.
-
Multiplexing and Labeling Strategies: Techniques like isobaric tags for relative and absolute quantitation (iTRAQ) or tandem mass tags (TMT) allow for the simultaneous identification and quantification of proteins from multiple samples in a single MS run, effectively increasing throughput.
-
Bioinformatics and Data Analysis: Robust bioinformatics pipelines and computational resources are essential for processing, interpreting, and managing the vast amounts of data generated by high-throughput protein sequencing experiments. Machine learning and artificial intelligence are increasingly being employed to improve data analysis and accelerate sequence determination.
High throughput protein sequencing is critical for large-scale proteomics studies, enabling the comparison of proteomes under different conditions (e.g., disease vs. healthy, drug-treated vs. untreated) and the discovery of novel protein biomarkers and drug targets.
Application of Protein Sequencing
Biopharmaceutical Characterization
-
Monoclonal Antibodies (mAbs) and Biologics: Ensuring the correct sequence, identifying sequence variants, and characterizing post-translational modifications (e.g., glycosylation, deamidation, oxidation) that impact efficacy, safety, and stability.
-
Biosimilar Development: Demonstrating structural similarity between a biosimilar and its reference product.
-
Therapeutic Peptides and Proteins: Confirming the sequence of synthesized peptides and recombinant proteins.
Biomarker Discovery and Validation
-
Identifying novel protein biomarkers in biofluids (blood, urine, CSF) for early disease detection, prognosis, and monitoring therapeutic responses in diseases like cancer, neurodegenerative disorders, and infectious diseases.
-
Validating candidate biomarkers identified through other proteomic approaches.
Drug Target Identification
-
Sequencing proteins from disease tissues or cells to identify novel therapeutic targets.
-
Understanding the mechanism of action of drugs by characterizing their protein interactions.
Fundamental Research
-
Understanding Protein Function: Elucidating how specific amino acid sequences and modifications influence protein activity, interactions, and cellular roles.
-
Proteomics: Large-scale protein identification, quantification, and characterization within complex biological systems, often coupled with differential expression analysis.
-
Structural Biology: Providing sequence information necessary for X-ray crystallography, NMR, and cryo-EM studies to determine 3D protein structures.
Enzyme Engineering and Protein Design
-
Verifying the sequence of engineered enzymes or proteins with altered properties.
-
Designing and validating novel peptides or proteins with specific functions.
Food Science and Agriculture
-
Identifying allergens in food products.
-
Detecting food adulteration.
-
Characterizing proteins in crops and livestock for improved nutritional value or disease resistance.
Forensics and Toxicology
-
Identifying protein components in forensic samples.
-
Analyzing protein toxins.
From unraveling the fundamental building blocks of life to driving the development of innovative therapies, the impact of protein sequencing is profound and ever-expanding. At Creative Biolabs, we combine decades of experience with cutting-edge technologies and a dedicated team of experts to provide comprehensive de novo sequencing services. We offer de novo antibody sequencing and de novo protein sequencing services, powered by our propriety DASS (Database Assisted Shotgun Sequencing) technology to meet the diverse protein research needs of our clients, driving innovation and advancement in the field of biomedical science.
Learn more about Creative Biolabs' de novo antibody sequencing services:
Reference
-
Smith, Matthew Beauregard, et al. "Estimating error rates for single molecule protein sequencing experiments." PLOS Computational Biology 20.7 (2024): e1012258. Distributed under Open Access license CC BY 4.0, without modification. https://doi.org/10.1371/journal.pcbi.1012258
All listed services and products are For Research Use Only. Do Not use in any diagnostic or therapeutic applications.