Accurate Peptide Sequencing: Principles, Techniques & Challenges
Introduction Fundamentals of Peptides Edman Sequencing MS Sequencing Key MS Approaches Challenges
Introduction to Peptides and Proteins
Peptides and proteins serve as essential components in molecular biology while carrying out numerous vital functions necessary for life. Deciphering their structure stands as a crucial step because the linear amino acid sequence provides the basic knowledge layer.
What are peptides and polypeptides?
Amino acids are the building blocks. Connecting amino acids through peptide bonds creates chains.
-
Peptides: Peptides consist of amino acid chains that usually have fewer than 50 residues in length. Amino acid chains function as hormones, neurotransmitters, growth factors and can become powerful toxins.
-
Polypeptides: Polypeptides represent extended sequences of amino acids without branches.
-
Proteins: Proteins are composed of one or more polypeptides which fold into intricate three-dimensional structures that determine their biological functions. The terms "polypeptide" and "protein" are sometimes considered synonymous but usually "protein" denotes an active molecule with complex folding.
Size differences and structural complexity primarily distinguish these molecules. The peptide bond represents the primary connection between molecules while the amino acid sequence serves as their defining feature.
Peptide bonds: structure and identification
An amide linkage known as the peptide bond develops between the carboxyl group (-COOH) of one amino acid and the amino group (-NH₂) of the subsequent amino acid. During this condensation reaction water is released as a byproduct.
The backbone structure (-N-Cα-C-) maintains planarity and rigidity from its partial double bond nature which affects polypeptide folding patterns. Peptide sequencing aims to discover the bonds between amino acids and determine the order of amino acids along the chain.
Fig. 1 Peptide structure.
The importance of amino acid sequences
The primary structure dictates:
-
Higher-Order Structure: A polypeptide's sequence dictates its folding into its functional three-dimensional shape.
-
Function: Biological activity depends on the organization of amino acids which form crucial active sites and binding domains along with necessary structural elements.
-
Protein Identification: Identifying proteins in complex biological samples depends upon sequencing distinct peptide fragments.
-
Drug Development: Analyzing the sequence of therapeutic peptides and protein targets is essential for their design processes as well as efficacy evaluations and quality control assessments.
-
Evolutionary Relationships: Comparing sequences across species reveals phylogenetic connections.
Fundamentals of Peptides
A deeper dive into the components and structure is essential before exploring sequencing methodologies.
Amino Acids: The Core Components
Twenty standard amino acids are encoded by the genetic code, each possessing a central carbon atom bonded to:
-
An amino group (-NH₂)
-
A carboxyl group (-COOH)
-
A hydrogen atom (-H)
-
A unique side chain (R-group)
The R-group determines the amino acid's specific chemical properties (e.g., size, charge, hydrophobicity, reactivity), influencing the peptide's overall characteristics.
Formation of Peptide Bonds
As mentioned, peptide bond formation is a dehydration synthesis reaction catalyzed by ribosomes during protein biosynthesis. In synthetic peptide chemistry, chemical coupling agents facilitate this process.
Understanding Peptide and Polypeptide Chains
Chains are directional, possessing:
-
An N-terminus: The end with a free amino group (-NH₂).
-
A C-terminus: The end with a free carboxyl group (-COOH).
By convention, peptide sequences are written from the N-terminus to the C-terminus.
Primary Structure: The Amino Acid Sequence
This linear arrangement (e.g., Ala-Gly-Ser-Met...) is the primary structure. Determining this exact order is the objective of peptide sequencing.
Classical Sequencing: Edman Degradation
Developed by Pehr Edman in the 1950s, Edman degradation was the cornerstone of protein sequencing for decades. It provides direct sequence information from the N-terminus.
The Step-by-Step Process of Edman Sequencing
Edman degradation is a cyclical chemical process:
-
Coupling: The peptide is reacted with phenyl isothiocyanate (PITC) under mildly alkaline conditions. PITC specifically targets the free N-terminal amino group, forming a phenylthiocarbamyl (PTC) peptide derivative.
-
Cleavage: The sample is treated with anhydrous acid (e.g., trifluoroacetic acid - TFA). This selectively cleaves the peptide bond between the first and second amino acid residues, releasing the N-terminal amino acid as an anilinothiazolinone (ATZ) derivative, leaving the rest of the peptide chain intact but one residue shorter.
-
Conversion: Aqueous acid treatment transforms the unstable ATZ derivative into the more stable phenylthiohydantoin (PTH) amino acid derivative.
-
Identification: Identification of the specific PTH-amino acid generally involves using High-Performance Liquid Chromatography (HPLC) to compare its retention time against established standards.
-
Cycle Repetition: The degradation cycle restarts with the shortened peptide that has gained a new N-terminal residue.
Identifying Amino Acids Sequentially (PTH-amino acids)
Each of the 20 standard amino acids yields a unique PTH derivative with a characteristic retention time on a specific HPLC column. By running known PTH-amino acid standards, a chromatogram library is created. The PTH-amino acid released in each cycle is identified by matching its peak's retention time to these standards.
Strengths and Limitations of Edman Sequencing
|
Feature
|
Strengths
|
Limitations
|
|
Methodology
|
Direct sequencing from N-terminus
|
Sequential, relatively slow (typically ~1 hour per cycle)
|
|
Sample Req.
|
Can work with relatively pure samples
|
Requires purified peptide/protein; μg to mg quantities often needed
|
|
Peptide Length
|
Reliable for ~30-50 residues per run
|
Incomplete reactions/side reactions accumulate, limiting read length
|
|
N-Terminus
|
Excellent for confirming N-terminal sequence
|
Fails if N-terminus is chemically blocked (e.g., acetylation, pyroglutamate)
|
|
Modifications
|
Poorly suited for identifying most PTMs
|
Most modifications interfere with the chemistry or identification
|
|
Mixtures
|
Generally unsuitable for complex mixtures
|
Requires single, pure sequence for clear results
|
|
Automation
|
Automated sequencers are available
|
Still requires significant hands-on time for setup and analysis
|
While powerful, Edman degradation has been largely superseded by mass spectrometry for high-throughput and complex analyses, though it remains valuable for specific applications like N-terminal confirmation of recombinant proteins.
Modern Sequencing: Mass Spectrometry (MS)
Mass spectrometry has revolutionized proteomics and peptide sequencing, offering unparalleled sensitivity, speed, and the ability to analyze complex mixtures and post-translationally modified peptides.
Why MS Revolutionized Peptide Sequencing
-
Sensitivity: MS can detect peptides at femtomole to attomole levels, orders of magnitude lower than Edman degradation.
-
Speed: Analysis of multiple peptides can occur within minutes to hours, enabling high-throughput studies (proteomics).
-
Versatility: Capable of handling complex mixtures (e.g., total cell lysates after protein digestion).
-
PTM Analysis: Can readily identify and locate various post-translational modifications (PTMs) based on mass shifts.
-
Blocked N-termini: Not an impediment, as sequencing relies on fragmentation patterns, not N-terminal chemistry.
Basic Principles of Mass Spectrometry: Ionization, Mass Analysis, Detection
A mass spectrometer measures the mass-to-charge ratio (m/z) of ions. The basic workflow involves:
-
Ionization: Converting analyte molecules (peptides) into gas-phase ions.
-
Mass Analysis: Separating the ions based on their m/z ratio using electric and/or magnetic fields within a mass analyzer.
-
Detection: Measuring the abundance of ions at each m/z value, generating a mass spectrum (Intensity vs. m/z).
Key MS Approaches for Sequencing
Tandem Mass Spectrometry (MS/MS): Fragmentation Analysis
This is the workhorse of MS-based sequencing.
-
MS1 Scan: A standard mass spectrum is acquired to detect peptide precursor ions present in the sample.
-
Precursor Selection: One specific precursor ion (based on its m/z) is isolated.
-
Fragmentation: The selected precursor ion is fragmented, typically by collision-induced dissociation (CID) or higher-energy collisional dissociation (HCD), where the ion collides with inert gas molecules. This breaks the peptide backbone primarily at the peptide bonds.
-
MS2 Scan: A mass spectrum of the resulting fragment ions is acquired.
Fragmentation typically produces b-ions (containing the N-terminus) and y-ions (containing the C-terminus). The mass difference between consecutive ions in a b-series or y-series corresponds to the mass of a specific amino acid residue. By analyzing the pattern of fragment ions (the MS/MS spectrum), the amino acid sequence can be deduced.
Liquid Chromatography-Mass Spectrometry (LC-MS / LC-MS/MS)
Complex samples such as tryptic digests of whole proteomes undergo initial peptide separation through liquid chromatography, particularly Reversed-Phase HPLC (RP-HPLC), before mass spectrometric analysis.
-
LC Separation: Reversed-Phase HPLC separates peptides based on their hydrophobic properties. The separation process simplifies the sample complexity that reaches the mass spectrometer during analysis.
-
Online Coupling: The ESI source receives the LC eluent through direct flow.
-
Data-Dependent Acquisition (DDA): The mass spectrometer quickly switches between MS1 scans which detect eluting peptides and MS/MS scans that fragment and sequence the most intense peptides from the MS1 scans.
LC-MS/MS technology enables researchers to identify and sequence numerous peptides from complicated mixtures through one experiment.
De Novo Sequencing: Building Sequences from Scratch
De novo sequencing directly identifies peptide sequences from MS/MS spectra without the need for a reference sequence database.
-
Process: The order of amino acids in peptides gets deduced through algorithms that analyze the mass differences between fragment ion peaks such as b and y in the MS/MS spectrum.
-
Application: Essential for sequencing novel peptides, antibodies (especially variable regions), peptides from organisms with unsequenced genomes, or identifying unexpected modifications.
-
Challenge: Requires high-quality, information-rich MS/MS spectra. Can be computationally intensive and may have ambiguities (e.g., Leu/Ile). However, Creative Biolabs has unprecedentedly achieved the precise identification of leucine and isoleucine by mass spectrometry, advancing breakthrough science on protein/antibody sequencing.
Fig. 2 Complete de novo sequencing of commercial monoclonal antibodies by deep learning models.1
Challenges in Peptide Sequencing
Peptide bond identification
While MS/MS fragmentation is effective, achieving complete fragmentation across every peptide bond (full b- and y-ion series) is rare. Gaps in the series can make definitive sequencing difficult, especially for de novo efforts. Proline residues, due to their cyclic structure, can lead to unique fragmentation patterns or suppress fragmentation altogether.
Identifying peptide sequence with ambiguity
-
Isobaric Residues: Leucine (Leu) and Isoleucine (Ile) have identical nominal and average masses. Distinguishing them often requires specific fragment ions (w- or d-ions generated by specific fragmentation methods) or high-resolution MS measurements of fragment ions, but they often remain ambiguous ('Xle' or L/I). Lysine (Lys) and Glutamine (Gln) are also isobaric in nominal mass but easily distinguished by accurate mass MS or their fragmentation behavior.
-
Post-Translational Modifications (PTMs): Modifications add mass and complexity. Identifying both the type and location of a PTM requires careful analysis of mass shifts in precursor and fragment ions. Some PTMs are labile and may be lost during ionization/fragmentation.
-
Incomplete Fragmentation: As mentioned, lack of key fragment ions can leave gaps or uncertainties in the sequence.
-
Sequence Rearrangements: Certain sequences can undergo gas-phase rearrangements during MS/MS, complicating spectral interpretation.
-
Low Abundance Peptides: Detecting and obtaining high-quality MS/MS spectra for low-abundance peptides in complex mixtures remains challenging.
Peptide and protein amino acid sequencing serves as a fundamental element in both biological research and biopharmaceutical development today. Creative Biolabs utilizes decades of experience with state-of-the-art equipment and advanced LC-MS/MS systems to produce precise peptide sequencing data. Our commitment enables us to deliver the essential sequence information required by our clients to speed up their research breakthroughs. Meanwhile, we offer de novo antibody sequencing and de novo protein sequencing services, powered by our propriety DASS (Database Assisted Shotgun Sequencing) technology to meet the diverse protein research needs of our clients, driving innovation and advancement in the field of biomedical science.
Learn more about Creative Biolabs' de novo antibody sequencing services:
Reference
-
Gueto-Tettay, Carlos, et al. "Multienzyme deep learning models improve peptide de novo sequencing by mass spectrometry proteomics." PLoS Computational Biology 19.1 (2023): e1010457. Distributed under Open Access license CC BY 4.0, without modification. https://doi.org/10.1371/journal.pcbi.1010457
All listed services and products are For Research Use Only. Do Not use in any diagnostic or therapeutic applications.