$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
RNA splicing is a crucial process that joins exons for translation and removes introns, facilitating RNA export and maintaining nucleic acid homeostasis. This tightly regulated mechanism operates in a temporal and spatial manner, contributing to transcriptome diversity and complexity1. Splicing is guided by key signals, including the 5' splice site (5'ss), branch site, and 3' splice site (3'ss), along with additional regulatory elements such as the polypyrimidine tract downstream of the branch site and the AG dinucleotide exclusion zone, which aid in 3'ss recognition2. Mechanistically, U1 small nuclear RNA (snRNA) pairs with the 5'ss, while the branch site interacts with U2 snRNA3. The coordinated activity of the branch site, polypyrimidine tract, and 3'ss facilitates the binding of U2 small nucleoproteins (snRNPs) and U2 auxiliary factors, stabilizing the spliceosome and positioning the branchpoint for nucleophilic attack on the 5'ss, thereby initiating splicing.
With the rapid advancement of sequencing technologies, the cost of whole-genome sequencing continues to decline, leading to an ever-expanding catalog of human genetic variants. It is estimated that 10-30% of disease-associated mutations in rare genetic disorders affect RNA splicing4,5,6, often producing aberrant gene products that may serve as therapeutic targets. However, evaluating the functional impact of intronic variants remains challenging due to the complexity of splicing signals, which are often redundant and degenerate. While the 5' and 3' splice sites are relatively well-characterized, branch sites, polypyrimidine tracts, and other splicing regulatory elements exhibit considerable sequence and positional variability in higher eukaryotes. Large-scale mapping studies have further demonstrated that multiple branch sites can exist within a single intron7,8,9,10, complicating the interpretation of intron-exon boundary variations.
Deep learning has been employed to assess how primary sequences contribute to splice site recognition4,11,12, revealing that splicing variants cluster at canonical splice sites while extending sparsely into the exon and the 3' region of introns. This pattern aligns with the established understanding that 5' splice site selection is primarily dictated by consensus sequences, whereas 3' splice site recognition depends on additional intronic elements, such as branch sites and polypyrimidine tracts. However, existing models have been trained to distinguish constitutive splice sites from alternative or artificial ones rather than focusing on intronic variants specifically. As a result, these computational tools show only moderate predictive accuracy, primarily identifying splice sites and exonic splicing variants13,14,15,16. In addition to predictive models, an experimental system capable of validating splicing variants in bulk would significantly enhance the identification and characterization of splicing defects.
Comprehensive RNA sequencing data linking disease-associated intronic variants to splicing phenotypes remain scarce due to their low frequency and the difficulty of predicting splicing outcomes from existing datasets. To address this gap, high-throughput splicing assays and computational models have been developed to systematically analyze splicing variants. Massively parallel splicing reporter assays (MaPSy) have been designed to evaluate the impact of variable sequences on splice-site selection. By incorporating sequence variations near the 5' and 3' splice sites or spanning entire intron-exon regions within fixed minigene backbones, MaPSy enables a functional assessment of splicing alterations. However, due to the limitations of bulk oligonucleotide synthesis, deep intronic variants and pseudo-exon activation are not captured in this approach.
The robustness of MaPSy has been validated using 70 independent splicing minigenes, yielding a Pearson's correlation of 0.8917. Notably, approximately 90% of splicing donor variants (+1 and +2) exhibited splicing defects, highlighting the assay's accuracy. Additionally, MaPSy with randomized branch site sequences has uncovered the degenerate nature of branch site recognition and its reliance on U2 core proteins18. Furthermore, a split-GFP MaPSy design coupled with fluorescence-activated cell sorting (FACS) has been used to investigate exon-skipping events caused by genetic variations. This approach revealed that 54% of splicing-disrupting variants reside in intronic regions, including canonical splice sites13, underscoring the significant role of intronic elements in splicing regulation. Collectively, these findings reinforce the importance of intronic sequences in splicing control and demonstrate the utility of MaPSy in identifying disease-associated splicing defects (Figure 1).

Figure 1: Experimental design of massively parallel splicing assay (MaPSy) of near-exon intronic mutations. Variants documented in databases of human disease were collected and synthesized as 5,307 pairs of oligos. Each oligo pair contains a reference and a variant allele across the 78-nucleotide (nt) intronic and 35-nt exonic regions. The oligos are flanked by common priming sites for amplification and ligation into 3-exon splicing minigenes. Accordingly, the synthesized region comprises the 3'ss of the second exon of the minigene. After minigene assembly, the pooled minigenes were spliced into human embryonic kidney (HEK293T) cells. The resulting spliced isoforms were harvested and resolved by amplicon sequencing. This figure has been adapted with permission from Chiang et al.17. Please click here to view a larger version of this figure.