$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Arginine (R)-methylation is a post translational modification (PTM) that decorates around 1% of the mammalian proteome1. Protein Arginine Methyltransferases (PRMTs) are the enzymes catalyzing R-methylation reaction by the deposition of one or two methyl groups to the nitrogen (N) atoms of the guanidino group of the side chain of R in a symmetric or asymmetric manner. In mammals, PRMTs can be grouped into three classes-type I, type II, and type III-depending on their capability to deposit both mono-methylation (MMA) and asymmetric di-methylation (ADMA), MMA and symmetric di-methylation (SDMA) or only MMA, respectively2,3. PRMTs mainly target R residues located within glycine- and arginine-rich regions, known as GAR motifs, but some PRMTs, such as PRMT5 and CARM1, can methylate proline-glycine-methionine-rich (PGM) motifs4. R-methylation has emerged as a protein modulator of several biological processes, such as RNA splicing5, DNA repair6, miRNA biogenesis7, and translation2, fostering the research on this PTM.
Mass Spectrometry (MS) is recognized as the most effective technology to systematically study global R-methylation at protein-, peptide-, and site-resolution. However, this PTM requires some particular precautions for its high-confidence identification by MS. First, R-methylation is substoichiometric, with the unmodified form of the peptides being much more abundant than the modified ones, so that mass spectrometers operating in the Data Dependent Acquisition (DDA) mode will fragment high-intensity unmodified peptides more often than their lower-intensity methylated counterparts8. Moreover, most MS-based workflows for R-methylated site identification suffer from limitations at the bioinformatic analysis level. Indeed, the computational identification of methyl-peptides is prone to high False Discovery Rates (FDR), because this PTM is isobaric to various amino acid substitutions (e.g., glycine into alanine) and chemical modification, such as methyl-esterification of aspartate and glutamate9. Hence, methods based on the isotope labeling of methyl groups, such as Heavy Methyl Stable Isotope Labeling with Amino Acids in Cell culture (hmSILAC), have been implemented as orthogonal strategies for confident MS-identification of in vivo methylations, significantly reducing the rate of false positive annotations10.
Recently, various proteome-wide protocols to study R-methylated proteins have been optimized. The development of antibody-based strategies for the immuno-affinity enrichment of R-methyl-peptides has led to the annotation of several hundreds of R-methylated sites in human cells11,12. Furthermore, many studies3,13 reported that coupling antibody-based enrichment with peptide separation techniques such as HpH-RP chromatography fractionation can boost the overall number of methyl-peptides identified.
This article describes an experimental strategy designed for the systematic and high-confidence identification of R-methylated sites in human cells, based on various biochemical and analytical steps: protein extraction from hmSILAC-labeled cells, parallel double enzymatic digestion with Trypsin and LysargiNase proteases, followed by HpH-RP chromatographic fractionation of digested peptides, coupled with antibody-based immuno-affinity enrichment of MMA-, SDMA-, and ADMA-containing peptides. All affinity-enriched peptides are then analyzed by high-resolution Liquid Chromatography (LC)-MS/MS in DDA mode, and raw MS data are processed by MaxQuant algorithm for identificationof R-methyl-peptides. Finally, the MaxQuant output results are processed with hmSEEKER, an in-house developed bioinformatics tool to search pairs of heavy and light methyl-peptides. Briefly, hmSEEKER reads and filters methyl-peptides identifications from the msms file, then matches each methyl-peptide to its corresponding MS1 peak in the allPeptides file, and, finally, searches the peak of the heavy/light peptide counterpart. For each putative heavy-light pair, the Log2 H/L ratio (LogRatio), Retention Time difference (dRT), and Mass Error (ME) parameters are calculated, and doublets that are located within user-defined cut-offs are labeled as true positives. The workflow of the biochemical protocol is described in Figure 1.