De Novo Protein Sequencing: A Comprehensive Guide by Mass Spectrometry
Introduction De Novo Protein Sequencing By Mass Spectrometry Challenges Applications
Introduction to Protein Sequencing
Overview of Protein Sequencing
The variety of essential functions for life relies on the activities of proteins. Protein function depends on specific three-dimensional structures which originate from the linear arrangement of amino acids known as the primary structure. Protein sequencing establishes the exact sequence of amino acids present in a protein. The knowledge of a protein's sequence forms the basis for understanding its identity along with its structure and function and it informs our understanding of its evolutionary history and possible alterations.
Importance of Protein Sequencing in Biology and Biotechnology
-
Functional Proteomics: We study protein functions through sequence analysis while focusing on enzymatic action comprehension and identification of protein interaction domains.
-
Structural Biology: The primary sequence of a protein serves as essential data for analyzing X-ray crystallography or cryo-EM results to build 3D structural models.
-
Biopharmaceutical Development: The characterization of therapeutic proteins such as monoclonal antibodies (mAbs) focuses on batch-to-batch consistency and detection of sequence variants or modifications.
-
Evolutionary Biology: Researchers analyze protein sequences from different species to trace evolutionary connections and examine protein preservation.
-
Systems Biology: Analyzing proteomic data together with genomic and transcriptomic data creates a comprehensive understanding of cellular processes.
Traditional Methods vs. De Novo Sequencing
The classic method for protein sequencing is Edman degradation, developed by Pehr Edman in the 1950s. This chemical method sequentially removes amino acids from the N-terminus of a peptide, which are then identified.
Pros: High accuracy for shorter peptides, directly identifies N-terminal residues.
Cons: Slow, requires relatively large amounts of pure protein, limited read length (typically < 50 residues), struggles with modified N-termini and complex mixtures.
Modern protein sequencing predominantly relies on Mass Spectrometry (MS). MS-based approaches can be broadly categorized:
-
Database-Dependent Sequencing
Matches experimental MS/MS spectra (fragment ion patterns) against theoretical spectra generated from known protein sequences stored in databases. This is highly effective when the organism's genome/proteome is well-characterized.
Interprets MS/MS spectra directly to deduce the peptide sequence without reference to any database. This is essential when dealing with unknown proteins or organisms lacking sequenced genomes.
What is De Novo Protein Sequencing?
Definition and Explanation of De Novo Sequencing
The term "de novo" is Latin for "from the beginning" or "anew." In the context of protein sequencing, de novo sequencing refers to the determination of a peptide's amino acid sequence directly from its tandem mass spectrometry (MS/MS) fragmentation data, without comparing it to a sequence database. De novo sequencing algorithms analyze the mass differences between fragment ions in an MS/MS spectrum to infer the sequence of amino acids that produced those fragments.
Why De Novo Sequencing is Needed
Database searching works well for recognized proteins but de novo sequencing becomes essential under multiple conditions.
-
Novel Proteins: Researchers need to identify proteins that current databases do not contain, such as those originating from unsequenced organisms or engineered proteins.
-
Antibody Sequencing: Researchers must identify the precise sequence arrangement of monoclonal and polyclonal antibodies since their variable regions (CDRs) essential for binding antigens display hypervariability beyond current database representations.
-
Venom/Toxin Research: The study of animal venoms reveals new toxins that feature distinct sequences and structural modifications.
-
Uncharacterized PTMs: Detecting peptides that exhibit unknown post-translational modifications beyond the scope of current database search algorithms.
-
Sequence Verification: Researchers need to validate recombinant protein sequences and discover any aberrant sequence variants.
-
Immunopeptidomics: The process of sequencing MHC-bound peptides on cell surfaces reveals peptides from non-canonical areas or spliced peptides that defy easy prediction systems.
Key Differences from Database-Dependent Sequencing
The choice between de novo and database-dependent methods hinges on the research question and sample origin.
|
Feature
|
Database-Dependent Sequencing
|
De Novo Sequencing
|
|
Requirement
|
Sequence Database (FASTA format)
|
High-quality MS/MS Spectra
|
|
Principle
|
Spectrum Matching (Experimental vs. Theory)
|
Spectrum Interpretation (Mass Differences)
|
|
Output
|
Peptide-Spectrum Match (PSM) Score, Sequence
|
Proposed Peptide Sequence(s), Confidence Score
|
|
Handling Unknowns
|
Limited (only identifies known sequences)
|
Primary Strength (identifies any sequence)
|
|
PTM Handling
|
Requires pre-specification of potential PTMs
|
Can potentially identify unexpected PTMs
|
|
Main Challenge
|
Database completeness, PTM complexity
|
Spectral quality, interpretation ambiguity
|
|
Computational Need
|
Generally lower per spectrum
|
Often higher, complex algorithms
|
De Novo Protein Sequencing by Mass Spectrometry
Mass Spectrometry in Protein Sequencing
Tandem mass spectrometry (MS/MS) remains the primary technique used in modern de novo sequencing. In a typical workflow:
-
Ionization: Proteins and peptides from enzymatic digestion (trypsin being an example) undergo ionization through Electrospray Ionization (ESI) or Matrix-Assisted Laser Desorption/Ionization (MALDI).
-
First Mass Analysis (MS1): Scientists measure the mass-to-charge ratios (m/z) of peptide ions that remain intact (precursor ions).
-
Isolation: Specific precursor ions of interest are selected.
-
Fragmentation: Selected precursor ions undergo fragmentation through techniques such as Collision-Induced Dissociation (CID), Higher-energy Collisional Dissociation (HCD), or Electron Transfer Dissociation (ETD). This fragmentation process generally occurs along the peptide backbone. CID/HCD: The fragmentation process produces b-ions that contain the N-terminus and y-ions that hold the C-terminus. ETD: The ETD fragmentation method produces c-ions and z-ions while maintaining the integrity of labile PTMs.
-
Second Mass Analysis (MS2): Fragment ion m/z ratios measurement results in the generation of the MS/MS spectrum.
Fig.1 The process of mass-spectrum generation.1
De novo sequencing relies on analyzing MS/MS spectra because the mass differences between peaks correspond to amino acid residue masses.
De Novo Sequencing Workflows Using Mass Spectrometry
-
Sample Preparation: The target protein or protein mixture requires isolation followed by purification.
-
Digestion: Specific proteases cleave proteins into smaller peptides such as trypsin which targets Lysine (K) and Arginine (R). To assist in assembly processes researchers may employ several enzymes when dealing with overlapping peptide sequences.
-
LC-MS/MS: The liquid chromatography (LC) system separates peptides which are then directly introduced into the mass spectrometer for MS1 scans and MS/MS fragmentation after precursor selection.
-
De Novo Interpretation: Sophisticated algorithms analyze the MS/MS spectra. They identify series of fragment ions (e.g., b- or y-ions) where consecutive peaks differ by the mass of an amino acid residue.
-
Peptide Assembly: Deduced peptide sequences are assembled, often using overlapping segments obtained from different enzymatic digests, to reconstruct the full protein sequence.
-
Validation: Manual inspection of spectra, comparison with orthogonal data (e.g., Edman sequencing of the N-terminus), and assessment of sequence coverage are crucial validation steps.
Challenges in De Novo Sequencing by Mass Spectrometry
Despite its power, de novo sequencing faces significant hurdles:
Error Accumulation
-
Incomplete Fragmentation: Not all peptide bonds break during fragmentation, leading to gaps in the fragment ion series within the spectrum.
-
Ambiguous Mass Gaps: Some amino acid combinations have very similar or identical masses (e.g., Leucine (L) vs. Isoleucine (I); Lysine (K) vs. Glutamine (Q) differ by only ~0.036 Da). High-resolution MS helps but doesn't always resolve ambiguity. PTMs add another layer of mass complexity.
-
Noisy Spectra: Low-intensity peaks, chemical noise, and co-fragmentation of multiple peptides can obscure the true signal.
-
Propagation of Errors: An incorrect residue assignment early in the interpretation can throw off the rest of the sequence derived from that spectrum.
Computational Complexity
-
Combinatorial Problem: The number of possible peptide sequences for a given precursor mass is vast. Algorithms must efficiently explore this space.
-
Scoring Functions: Developing accurate scoring functions to rank the likelihood of different sequence candidates derived from a spectrum is challenging.
-
Assembly Algorithms: Assembling potentially error-prone peptide sequences into a full protein sequence requires sophisticated algorithms, similar to genome assembly but often with less coverage and more ambiguity.
-
PTM Identification: Incorporating unknown PTMs significantly increases the search space and complexity.
Applications of De Novo Protein Sequencing
The ability to sequence proteins without prior knowledge unlocks critical applications:
Drug Discovery
-
Antibody Sequencing: Antibody sequencing remains vital throughout therapeutic antibody development as well as characterization processes and biosimilar production while securing patents. De novo sequencing accurately identifies the variable (Fv) and constant (Fc) regions.
-
Target Characterization: Identifying and sequencing novel protein targets or understanding modifications on known targets.
-
Mechanism of Action Studies: Identifying protein interaction partners or drug-induced PTMs.
Proteomics Research
-
Characterizing Unknown Proteomes: We sequence proteins from organisms that have yet to have their genomes sequenced such as exotic species or environmental isolates.
-
Venomics/Toxinology: The field of Venomics and Toxinology involves discovering and analyzing new bioactive peptides and proteins found in venoms and toxins.
-
Analysis of Proteoforms: Identify and sequence protein variations that originate from post-translational modifications (PTMs), alternative splicing events or sequence variants.
-
Quality Control: Verifying the sequence of recombinant proteins produced for research or therapeutic use.
Biomarker Identification
-
Discovery of Novel Biomarkers: Identifying proteins or specific PTMs in complex biological fluids (plasma, urine, CSF) whose sequences might not be in standard databases or which represent disease-specific modifications.
-
Understanding Disease Mechanisms: Sequencing proteins involved in pathogenesis that may be novel or unexpectedly modified.
At Creative Biolabs, we offer de novo antibody sequencing and de novo protein sequencing services, powered by our propriety DASS (Database Assisted Shotgun Sequencing) technology to meet the diverse protein research needs of our clients, driving innovation and advancement in the field of biomedical science.
Learn more about Creative Biolabs' de novo antibody sequencing services:
Reference
-
Wu, Ruitao, et al. "Denovo-GCN: De novo peptide sequencing by graph convolutional neural networks." Applied Sciences 13.7 (2023): 4604. Distributed under Open Access license CC BY 4.0, without modification. https://doi.org/10.3390/app13074604
All listed services and products are For Research Use Only. Do Not use in any diagnostic or therapeutic applications.