$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
The identification of the complete set of coding elements in the genome has been a major goal since the initiation of the Human Genome Project, and remains a central objective toward the understanding of biological systems and the etiology of genetic-based diseases1,2,3,4. Advances in NGS techniques have led to the production of whole genome sequences for an extensive number of organisms, including vertebrates, invertebrates, yeast, and plants5. Additionally, high-throughput transcriptional sequencing methods have further revealed the complexity of the cellular transcriptome, and identified thousands of novel RNA molecules with both protein-coding and noncoding functions6,7. Decoding this vast amount of sequence information is an ongoing process, and challenges remain with comprehensive gene annotation efforts8.
The recent development of translational profiling methods, including ribosome profiling9,10 and poly-ribosome sequencing11, have provided evidence indicating that hundreds of noncanonical translation events map to currently unannotated sORFs throughout the genome, with the potential to generate small proteins called microproteins or micropeptides12,13,14,15,16,17. Microproteins have emerged as a novel class of versatile proteins previously overlooked by standard gene annotation methods due to their small size (<100 amino acids) and lack of classical protein-coding gene characteristics8,12,18,19,20. Microproteins have been described in virtually all organisms, including yeast21,22, flies17,23,24, and mammals25,26,27,28, and have been shown to play critical roles in diverse processes, including development, metabolism, and stress signaling19,20,29,30,31,32,33,34. Thus, it is imperative to continue to mine the genome for additional members of this long-overlooked class of functional small proteins.
Despite the widespread recognition of the biological importance of microproteins, this class of genes remains vastly underrepresented in genome annotations, and their accurate identification continues to be an ongoing challenge that has hindered progress in the field. Various computational tools and experimental methods have recently been developed to overcome the difficulties associated with identifying microprotein-coding sequences (discussed extensively in several comprehensive reviews8,35,36,37). Many recent microprotein identification studies38,39,40,41,42,43,44,45,46,47 have relied heavily on the use of one such algorithm called PhyloCSF48,49, a powerful comparative genomics approach that can be leveraged to distinguish conserved protein-coding regions of the genome from those that are noncoding.
PhyloCSF compares codon substitution frequencies (CSF) using multi-species nucleotide alignments and phylogenetic models to detect evolutionary signatures of protein-coding genes. This empirical model-based approach relies on the premise that proteins are primarily conserved at the amino acid level rather than the nucleotide sequence. Therefore, synonymous codon substitutions, which encode the same amino acid, or codon substitutions to amino acids with conserved properties (i.e., charge, hydrophobicity, polarity) are scored positively, while non-synonymous substitutions, including missense and nonsense substitutions, score negatively. PhyloCSF is trained on whole-genome data and has proven to be effective in scoring short portions of a coding sequence (CDS) in isolation from the full sequence, which is necessary when analyzing microproteins or individual exons of standard protein-coding genes48,49.
Notably, the recent integration of the PhyloCSF track hubs in the University of California Santa Cruz (UCSC) Genome Browser49,50,51 enables investigators of all backgrounds to easily access a user-friendly interface to query genomic regions of interest for protein-coding potential. The protocol outlined below provides detailed instruction on how to load the PhyloCSF track hubs on the UCSC Genome Browser and subsequently interrogate genomic regions of interest to probe for high-confidence protein-coding regions (or the lack thereof). Additionally, in the case where a positive PhyloCSF score is observed, steps are delineated to further analyze microprotein-coding potential and efficiently generate multiple species alignments of the identified amino acid sequences to illustrate cross-species sequence conservation. Lastly, several additional publicly available resources and tools are introduced in the discussion to survey identified microprotein characteristics, including predicted domain structures and insight into putative microprotein function.