Method Article

Massively Parallel Splicing Assay to Examine Splicing Errors Caused by Disease-Related Intronic Variants

DOI:

10.3791/68984

September 9th, 2025

 ,  ,  ,  ,  , 

Corresponding Authors: Chien-Ling Lin <mbcllin@gate.sinica.edu.tw>

* These authors contributed equally

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Here, we present a detailed protocol for performing massively parallel splicing assays (MaPSy), which employ minigene constructs to systematically evaluate intronic variants in bulk. This approach enables high-throughput analysis of variant-induced splicing changes in cells through amplicon sequencing, providing functional assessments of their impact on pre-mRNA splicing.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Splicing errors represent 10-30% of the pathogenic mutations responsible for rare genetic disorders. RNA splicing ensures proper gene expression by selectively joining exons and removing introns, with key regulatory sequences being located within the introns. The 5' splice site and branch site interact with small nuclear RNAs to form the spliceosome's recognition complex, while elements such as the polypyrimidine tract and splicing enhancers/silencers recruit proteins to regulate spliceosome assembly. Predicting splicing disruptions from intronic variants is challenging due to the complexity of these interactions.

Intronic variants, comprising 90% of natural human gene variations, may disrupt canonical splicing and give rise to disease. To investigate this possibility, we developed a massively parallel splicing assay (MaPSy) to assess patient-identified intronic variants. Synthesized oligonucleotides with reference or variant sequences were ligated into splicing minigenes containing promoter and polyadenylation signals. Each construct included two constant exons flanking a middle exon that harbored the variable intron-exon junction sequence of interest. The cellular splicing efficiency of the variant sequences was compared to reference counterparts, allowing us to identify significant disruptions as splicing variants.

The results of the MaPSy can be validated through additional approaches, such as minigene assays or CRISPR-mediated genome editing in vivo. Furthermore, aggregate analysis of the disrupted junctions can provide deeper insights into splicing mechanisms and the molecular basis of diseases associated with splicing errors.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

RNA splicing is a crucial process that joins exons for translation and removes introns, facilitating RNA export and maintaining nucleic acid homeostasis. This tightly regulated mechanism operates in a temporal and spatial manner, contributing to transcriptome diversity and complexity1. Splicing is guided by key signals, including the 5' splice site (5'ss), branch site, and 3' splice site (3'ss), along with additional regulatory elements such as the polypyrimidine tract downstream of the branch site and the AG dinucleotide exclusion zone, which aid in 3'ss recognition2. Mechanistically, U1 small nuclear RNA (snRNA) pairs with the 5'ss, while the branch site interacts with U2 snRNA3. The coordinated activity of the branch site, polypyrimidine tract, and 3'ss facilitates the binding of U2 small nucleoproteins (snRNPs) and U2 auxiliary factors, stabilizing the spliceosome and positioning the branchpoint for nucleophilic attack on the 5'ss, thereby initiating splicing.

With the rapid advancement of sequencing technologies, the cost of whole-genome sequencing continues to decline, leading to an ever-expanding catalog of human genetic variants. It is estimated that 10-30% of disease-associated mutations in rare genetic disorders affect RNA splicing4,5,6, often producing aberrant gene products that may serve as therapeutic targets. However, evaluating the functional impact of intronic variants remains challenging due to the complexity of splicing signals, which are often redundant and degenerate. While the 5' and 3' splice sites are relatively well-characterized, branch sites, polypyrimidine tracts, and other splicing regulatory elements exhibit considerable sequence and positional variability in higher eukaryotes. Large-scale mapping studies have further demonstrated that multiple branch sites can exist within a single intron7,8,9,10, complicating the interpretation of intron-exon boundary variations.

Deep learning has been employed to assess how primary sequences contribute to splice site recognition4,11,12, revealing that splicing variants cluster at canonical splice sites while extending sparsely into the exon and the 3' region of introns. This pattern aligns with the established understanding that 5' splice site selection is primarily dictated by consensus sequences, whereas 3' splice site recognition depends on additional intronic elements, such as branch sites and polypyrimidine tracts. However, existing models have been trained to distinguish constitutive splice sites from alternative or artificial ones rather than focusing on intronic variants specifically. As a result, these computational tools show only moderate predictive accuracy, primarily identifying splice sites and exonic splicing variants13,14,15,16. In addition to predictive models, an experimental system capable of validating splicing variants in bulk would significantly enhance the identification and characterization of splicing defects.

Comprehensive RNA sequencing data linking disease-associated intronic variants to splicing phenotypes remain scarce due to their low frequency and the difficulty of predicting splicing outcomes from existing datasets. To address this gap, high-throughput splicing assays and computational models have been developed to systematically analyze splicing variants. Massively parallel splicing reporter assays (MaPSy) have been designed to evaluate the impact of variable sequences on splice-site selection. By incorporating sequence variations near the 5' and 3' splice sites or spanning entire intron-exon regions within fixed minigene backbones, MaPSy enables a functional assessment of splicing alterations. However, due to the limitations of bulk oligonucleotide synthesis, deep intronic variants and pseudo-exon activation are not captured in this approach.

The robustness of MaPSy has been validated using 70 independent splicing minigenes, yielding a Pearson's correlation of 0.8917. Notably, approximately 90% of splicing donor variants (+1 and +2) exhibited splicing defects, highlighting the assay's accuracy. Additionally, MaPSy with randomized branch site sequences has uncovered the degenerate nature of branch site recognition and its reliance on U2 core proteins18. Furthermore, a split-GFP MaPSy design coupled with fluorescence-activated cell sorting (FACS) has been used to investigate exon-skipping events caused by genetic variations. This approach revealed that 54% of splicing-disrupting variants reside in intronic regions, including canonical splice sites13, underscoring the significant role of intronic elements in splicing regulation. Collectively, these findings reinforce the importance of intronic sequences in splicing control and demonstrate the utility of MaPSy in identifying disease-associated splicing defects (Figure 1).

Gene splicing diagram; oligo synthesis, PCR constructs, HEK293T expression, amplicon sequencing.
Figure 1: Experimental design of massively parallel splicing assay (MaPSy) of near-exon intronic mutations. Variants documented in databases of human disease were collected and synthesized as 5,307 pairs of oligos. Each oligo pair contains a reference and a variant allele across the 78-nucleotide (nt) intronic and 35-nt exonic regions. The oligos are flanked by common priming sites for amplification and ligation into 3-exon splicing minigenes. Accordingly, the synthesized region comprises the 3'ss of the second exon of the minigene. After minigene assembly, the pooled minigenes were spliced into human embryonic kidney (HEK293T) cells. The resulting spliced isoforms were harvested and resolved by amplicon sequencing. This figure has been adapted with permission from Chiang et al.17. Please click here to view a larger version of this figure.

Protocol

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

1. Synthesis of the MaPSy oligonucleotide (oligo) library

  1. Basic oligo structure: Design each oligo in the 155-nuncleotide (nt) pool to include a 35-nt exonic sequence and an 80-nt intronic sequence, forming a 115-nt gene-specific region flanked by 20-nt common priming sequences at each end (Figure 2A).
  2. Variant collection: To retain motif integrity and focus on regulatory regions affecting splicing, collect the intronic variants that are:
    -78 to +10 nt from the 3'ss
    -3 to +30 nt from the 5'ss
    NOTE: Clinical variants can be sourced from ClinVar19, low-frequency entries in dbSNP20, relevant publications, and additional clinical databases. Genome coordinates can be used to intersect variants with the desired genomic regions (e.g., using BEDTools21). Only single-nucleotide polymorphisms (SNPs) and small insertions/deletions (indels) under 15 nt should be selected to minimize size variability within the oligo pool for bulk synthesis.
  3. Basic minigene structure: To preserve the splicing junction structure, each minigene backbone contains three exons: two constant outer exons encoding fragments of EGFP (from pGint, Addgene Plasmid #24217) and a middle exon derived from exon 15 of human CAMTA2 (Figure 2B). Digest both the CAMTA2 fragment and pGint with BamHI and SalI restriction enzymes and ligate the CAMTA2 fragment into pGint between the EGFP exons, creating pGint-CAMTA2 three-exon minigenes (Supplementary File 1).
    NOTE: The middle exon, i.e., CAMTA2 exon 15, should exhibit an intermediate splicing efficiency when expressed in target cells. This characteristic is essential as it provides a balanced baseline, allowing for the detection of both increases and decreases in splicing efficiency due to variant effects.
  4. Priming site design: Append the insertion site sequence to both ends of the library so that each oligo contains ~20-nt flanking sequences that overlap with the intended insertion site within the middle exon/intron (Figure 2A).
    NOTE: If the oligo pool is designed for multiple minigene constructs, multiple priming sites can be incorporated.
  5. Barcode design: For intronic variants, exonic barcodes are essential to distinguish genotypes from spliced products, as intronic sequences are lost post-splicing. To avoid any splicing alterations due to barcode sequences (i.e., the barcode effect), position the barcodes distally within the exons, away from the splice sites (i.e., right next to the exonic priming site).
    NOTE: It is recommended to use multiple barcodes per variant where possible. The required barcode length depends on library complexity. For example:
    A 1-nt barcode is sufficient if a reference allele pairs with only one variant.
    A 2-nt barcode is preferable if a reference allele pairs with six different variants.
    Libraries with multiple variants closely related to a single reference allele (e.g., differing by only one nucleotide) may increase the complexity of the analysis.
  6. Ordering oligos: Order oligos in FASTA file format, with each variant junction paired alongside its reference allele.
    ​NOTE: For junctions with multiple variants, only one reference allele is needed per junction. Pooled sequence synthesis services are available from companies such as GeneScript (https://www.genscript.com/gentitan-oligo-pools.html), Twist Bioscience (https://www.twistbioscience.com/products/oligopools), IDT (https://sg.idtdna.com/pages/products/custom-dna-rna/dna-oligos/custom-dna-oligos/opools-oligo-pools), and Agilent (https://www.agilent.com/en/product/oligo-pools-oligo-gmp-manufacturing/pooled-oligo-synthesis). Synthesis length, library capacity, and costs may vary depending on the provider and region.

Oligo and MaPSy minigene design diagram for splice site analysis and variant integration.
Figure 2: Oligo and splicing minigene design for MaPSy. (A) Design for pooled oligo synthesis of the 3' or 5' end of introns. The diagrams illustrate the basic structure of the 155-nt oligos. The actual capacity of oligo synthesis depends on the company selected for production. (B) Design of the MaPSy minigenes. A three-exon minigene (ii) was modified from pGint plasmid (i), and the splice sites were replaced by the pooled oligos to introduce splicing signal variations (iii). Please click here to view a larger version of this figure.

2. Construction of MaPSy-library DNA templates

  1. Initial amplification upon receiving the oligo pool
    1. Upon receipt of the oligo pool, amplify 10-50 ng of the oligo library with the designed flanking priming sites by a 100 µL polymerase chain reaction (PCR) using High-Fidelity DNA Polymerase to convert the oligos into double strands (LibF and LibR primers in Table 1, thermocycler settings in Table 2).
    2. Clean up the PCR products using purification columns. In detail, bind the PCR product to the column membrane, then wash with buffers containing 70% ethanol to remove residual primers, salts, and polymerase. Finally, elute the purified PCR product using low-salt buffer or nuclease-free water.
    3. Verify the amplified oligo size by gel electrophoresis. Load 5 µL of the purified PCR product onto a 1.5% agarose gel prepared in 1× TAE buffer (40 mM Tris base, 20 mM acetic acid, 1 mM EDTA, pH 8.0).
      NOTE: Limit PCR amplification cycles to 15 to prevent overamplification and PCR bias. A secondary fuzzy band, typically larger than the target product, may indicate overamplification and incomplete annealing within the oligo pool. Save a portion of the PCR products and sequence the oligo pool to assess quality and error rates before proceeding (stop point).
  2. Construction of MaPSy splicing minigenes - Initial PCR.
    1. Backbone fragments: Use 0.5 ng of the pGint-CAMTA2 plasmid backbone as the source for the two main PCR fragments
      1. PCR Product 1: Amplify the CMV promoter, the first exon (N-terminal EGFP), and part of the first intron (Table 1 and Table 2).
      2. PCR Product 3: Amplify part of the middle exon (CAMTA2 exon 15), the second intron, the third exon (C-terminal EGFP), and the SV40 polyadenylation signal (Figure 3, Table 1, and Table 2).
    2. Library amplicon: PCR Product 2: For the 3'ss library, amplify the library amplicons with overlapping sequences at each end to match the inner ends of PCR Product 1 and PCR Product 3, allowing for efficient integration at the 3'ss of CAMTA2 exon 15 (Figure 3, Table 1, and Table 2).
      NOTE: For the 5'ss library, replace the 5' end of the middle exon junction with the library sequence.
  3. After PCR, clean up all products using purification columns as step 2.1.2. To prevent contamination from the template vector, purify the desired products from PCR1 and PCR3 by agarose gel extraction. Specifically, after electrophoresis, excise ~ 100 mg of agarose gel containing the target PCR product and dissolve it in the binding buffer. Bind the dissolved gel solution to the column membrane, then proceed with the standard PCR purification protocol.
  4. Construction of MaPSy splicing minigenes - Overlap extension PCR: Perform one or sequential rounds of overlapping PCR to ligate the three main fragments (PCR Product 1, the library amplicon, and PCR Product 3; use ~ 20 ng each as template) into a full-length splicing minigene (use the most outside primers CAMGFPF and CMVGFPR in Table 1, and thermocycler settings in Table 2).
    NOTE: The final product contains the CMV promoter and three exons with the first 3'ss derived from the oligonucleotide library, and it is transfection-ready for cell-based experiments.
  5. Following assembly, clean up the full-length DNA templates using PCR purification columns. If non-specific bands are observed in the gel, perform gel extraction to isolate precisely the desired full-length product as step 2.3.
    ​NOTE: Highly similar sequences in the library may lead to incomplete assembly. If full assembly in a single PCR reaction is difficult, sequential overlapping PCRs (starting with two fragments, then adding a third) may be attempted. To improve PCR specificity, consider adjusting the annealing temperature or implementing touch-down PCR to enhance the binding specificity for complex templates. Prepare sufficient MaPSy constructs to perform at least four independent experiments. Save a small aliquot of each construct for next-generation sequencing to verify sequence integrity, as some oligonucleotides may not amplify efficiently and result in incomplete assemblies (stop point).

DNA library construction using PCR; splicing minigenes diagram; shows primers, steps, and products.
Figure 3: Workflow of library construction. (A) Primers used in the overlapping PCR (see also Table 1). (B) The procedure for overlapping PCR. In brief, oligo pools and the other parts of the splicing minigenes were amplified by 25 PCR cycles. The fragment containing the promoter and the first exon (PCR product 1) was stitched to the oligo pool (PCR product 2) by overlapping PCR using 20 amplification cycles. Then, the stitched product (PCR product 1+2) was further stitched to the fragment containing the 3rd exon and polyadenylation signal (PCR product 3) using 20 amplification cycles to obtain the final construct (PCR product 1+2+3). This figure has been adapted with permission from Chiang et al.17. Please click here to view a larger version of this figure.

Construction of MaPSy-library primernote
CMVGFPFCCGCCATGCATTAGTTATTAATAG PCR product 1
LibR2CAGGTCTTCAGGCCCCAGCCPCR product 1
LibFGGCTGGGGCCTGAAGACCTGPCR product 2
LibRAAGGCGCACATGACCCCGGGPCR product 2
LibF2CCCGGGGTCATGTGCGCCTTPCR product 3
CMVGFPRGGACAAACCACAACTAGAATGCPCR product 3
MaPSy-library PCR primer for amplicon sequencing
P7-Lib0FGTCTCGTGGGCTCGGAGATGTGTATAAGAGACAGAAGTTCAGCGTGTCCGGCGA
P7-Lib1FGTCTCGTGGGCTCGGAGATGTGTATAAGAGACAGNAAGTTCAGCGTGTCCGGCGA
P7-Lib2FGTCTCGTGGGCTCGGAGATGTGTATAAGAGACAGNNAAGTTCAGCGTGTCCGGCGA
P7-Lib3FGTCTCGTGGGCTCGGAGATGTGTATAAGAGACAGNNNAAGTTCAGCGTGTCCGGCGA
P5-Lib0RITCGTCGGCAGCGTCAGATGTGTATAAGAGACAGCGAAGGCTCCTGTCTCTGTAGT
P5-Lib1RITCGTCGGCAGCGTCAGATGTGTATAAGAGACAGNCGAAGGCTCCTGTCTCTGTAGT
P5-Lib2RITCGTCGGCAGCGTCAGATGTGTATAAGAGACAGNNCGAAGGCTCCTGTCTCTGTAGT
P5-Lib3RITCGTCGGCAGCGTCAGATGTGTATAAGAGACAGNNNCGAAGGCTCCTGTCTCTGTAGT
Validation primer
Lib0F (Library validation)AAGTTCAGCGTGTCCGGCGA
Lib0R (Library validation)CAGCTTGCCGTAGGTGGCAT

Table 1: Primers used in this protocol.

PCR reaction mix
ComponentsVolume
ddH213.4 μL
5x HF Buffer4 μL
10 mM dNTPs 0.4 μL
Forward primer 10 µM0.5 μL
Reverse primer 10 µM0.5 μL
cDNA 1 μL
High–Fidelity DNA Polymerase 0.2 μL
Run the PCR program (20 µL reaction)
TempTimeCycles
Initial Denaturation 98°C 1 min 1
Denaturation 98°C 30 s 10–20
Annealing65°C 30 s 
Extension 72°C30 s Back to denaturation 
Final extension72°C 5 min 1
Hold 4°C Hold Hold 

Table 2: Thermocycler settings.

3. Expression and recovery of splicing minigenes from mammalian cells

  1. Cell culture: Maintain HEK293T cells in Dulbecco's Modified Eagle's Medium (DMEM) supplemented with 10% fetal bovine serum (FBS), 100 units/mL penicillin and streptomycin, and 2 mM L-glutamine in 5% CO2 at 37 °C.
  2. Transfection of splicing minigenes:
    1. Plate 5 × 105 HEK293T cells in 2 mL of medium in 6-well plates 24 h before transfection.
    2. Transfect cells with 1-2 µg of MaPSy constructs, pre-incubated for 5 min with 3.75 µL transfection reagent, following the manufacturer's protocol. Successful transfection is indicated by the presence of EGFP signals from the spliced library.
    3. Lyse the cells using 250 µL Trizol (or equivalent) 24 h post-transfection.
  3. RNA extraction: Extract total RNA using an RNA miniprep kit, following the manufacturer's protocol. Specifically, bind RNA in the Trizol to the column membranes, wash with buffers containing 70% ethanol to remove salts, proteins and other impurities, and elute the purified RNA using low-salt buffer or nuclease-free water.
    CAUTION: Trizol (or equivalent) is highly corrosive and toxic. Exposure can lead to serious chemical burns, permanent scarring, and kidney failure.
    NOTE: RNA can be stored in Trizol (or an equivalent Trizol-based reagent, such as TOOLSmart RNA extractor by TOOLS) at -80 °C for up to one year (stop point). Purified RNA can be stored at -80 °C for up to two years if freeze-thaw cycles are minimized (stop point).
  4. Reverse transcription: Prepare cDNA from 2 µg total RNA using reverse transcriptase with random hexamers, following the manufacturer's protocol. Specifically, incubate the RNA with 50 µM random hexamers at room temperature for 10 min to allow annealing, and then perform reverse transcription at 55 °C for 10 min.
    NOTE: cDNA can be stored at -20 °C for up to one year (stop point).
  5. Amplification of the spliced minigenes: Perform PCR using the specific priming sequences for the spliced minigenes (Table 2). Use the minimal number of cycles necessary to visualize the product on an agarose gel.
    NOTE: Mixed bands of unspliced and spliced products should be visible on the gel. Bands may appear diffuse due to the mixed population of DNA species (Figure 4A).
  6. Clean up the PCR product using purification columns as step 2.1.2.
  7. Attachment of sequencing adapters: Perform a final round of PCR to attach sequencing adaptor sequences to the amplicon ends. Include 0-3 random nucleotides at the end of the amplicons to ensure balanced fluorescence detection on the NextSeq platform (Table 1).
  8. Clean up the PCR product using purification columns as step 2.1.2.

4. Amplicon sequencing and analysis

  1. Short-read sequencing: Subject the PCR amplicons to 150 paired-end sequencing using Illumina Miseq, Novaseq, or equivalent, through a core facility or commercial service.
  2. Alignment:
    1. Create a reference genome: Create a synthetic "reference genome" by labeling each unique genotype as a separate chromosome. The reference genome includes the synthetic exons and introns contained within the amplicons.
    2. Align the sequencing reads: Align the paired-end reads to the reference genome using HISAT222,23, adjusting the parameters to control for splicing-specific elements (see Supplementary File 2 for command-line details).
    3. Select high-quality reads: Convert SAM to BAM format, filter for high-quality reads (mapping quality ≥60), and then sort and index the BAM files24 (see Supplementary File 2 for command-line details).
    4. Identify splice junctions and calculate the junction reads: Quantify the splice junction usage by extracting exon-skipping events from CIGAR strings in the aligned BAM file. Identify the junction-spanning reads based on "N" operations, and aggregate read counts per junction coordinate and strand to assess splicing patterns (see Supplementary File 2 for command-line details).
    5. Classify the canonical splice sites: Classify splice sites as canonical if they match the annotated GT-AG junctions.
      NOTE: In rare instances, reference alleles in MaPSy may utilize noncanonical splice sites. Reads that lack junctions spanning the designated splice site position are retained as unspliced reads.
  3. Statistical analysis to identify splicing variants: Categorize reads into three groups:
    (1) spliced versus unspliced reads;
    (2) canonical versus noncanonical among all spliced reads;
    (3) canonical versus noncanonical plus unspliced reads.
    Conduct a two-sided Fisher's exact test, followed by false discovery rate (FDR) correction (Table 3), to evaluate the impact of variants on splicing efficiency and accuracy.
  4. Filter for high-confidence splicing variants: Classify both reference/variant pairs exceeding 100 read counts with a q-value of less than 0.05 across four repeats as significant. Then, consider candidates with a 2-fold odds ratio change with either the reference or the variant allele having >5% unspliced and noncanonical reads, high-confidence splicing variants (Figure 4B).
ReadsSplicedUnspliced and/or Noncanonical
Referenceab
Variantc

Table 3: Two-by-two table for Fisher's exact test.

5. Validation

  1. Minigene splicing for validation:
    1. Oligo synthesis and amplification: Synthesize the DNA oligos of selected MaPSy candidate sequences individually (e.g., by Integrated DNA Technologies). Then, amplify the oligos using the designed flanking sequences into double strands by PCR using High-Fidelity DNA Polymerase (Table 2).
    2. Minigene cloning: Digest both the resulting PCR products and the pGint-CAMTA2 by BbsI and SmaI and ligate the digested product by DNA ligases.
    3. Transfection: Transfect the resulting constructs into HEK293T cells using a transfection reagent as step 3.2.2.
    4. RNA extraction: Extract RNA from the transfected cells, as described in step 3.3.
    5. Reverse transcription: Perform reverse transcription polymerase chain reaction (RT-PCR) using random hexamers.
    6. Splicing isoform amplification: Amplify the splicing isoforms with primers targeting the first two exons of the minigene (Lib0F and Lib0Rl, Table 1, thermocycler settings in Table 2). Resolve the amplified products by electrophoresis and visualize using the Gel Doc System.
    7. Quantification: Quantify the signal intensity of each splicing isoform using ImageJ (National Institutes of Health, USA)25,26. Alternatively, use the DNA Screening Kit using the eGENE HDA-GT12 high-performance nucleic acid analyzer to quantify the intensity and molecular weight of the PCR products.
    8. Isoform extraction and confirmation: Isolate each isoform by agarose gel extraction as step 2.3, and confirm the splicing outcome by Sanger sequencing of the PCR products through a core facility or commercial service, assessing for normal splicing, intron inclusion, and exon skipping.
  2. Multi-exon splicing minigenes:
    1. After validating the MaPSy results, select the variant of interest and clone three to five exons from genomic DNA (gDNA) to provide a more genomic context for the observed splicing defect.
    2. Perform site-directed mutagenesis using mutagenizing primers and overlapping PCR to assemble minigene constructs harboring the desired sequence alteration at the target site.
      NOTE: If the flanking introns are too long for cloning, retain approximately 300 nt of the intronic sequence for each splice site to ensure a proper splicing context.
    3. Perform the cell-based splicing assay as described above in 5.1.
  3. Cellular validation: To validate the splicing effect in cells, use template-based CRISPR editing to alter the sequence of the selected variants in an appropriate cellular model.
    NOTE: When possible, use human samples carrying the specific variant to assess the splicing pattern directly.

Results

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Following cellular splicing of MaPSy constructs, both spliced and unspliced products are present as a mixture. Due to the library's size diversity and potential for noncanonical splicing, both types of products may appear somewhat diffuse on a gel. In constructs targeting the 3' end, the second intron, containing partial adenoviral sequences, tends to splice very robustly (Figure 4A).

In MaPSy experiments, approximately 10-30% of disease-relevant variants exhibit altered splicing, with the majority of significant cases showing reduced canonical splicing efficiency (Figure 4B). Although noncanonical splicing may modestly shift the pool of splicing variants, including these events in the analysis is essential to identify variants that promote noncanonical splice sites within a splice site competition framework (Figure 4C). A prior study has indicated that selection of the 3'ss is particularly sensitive to competition from mutations in adjacent regions17.

Gel electrophoresis, data analysis, Venn diagram; DNA splicing outcomes, visualized experiment results.
Figure 4: Representative results of MaPSy on disease-relevant intronic variants. (A) RT-PCR analysis showing both spliced and unspliced products from minigenes containing the library of intronic variants after cellular splicing. The representative gel image shows the range of splicing outcomes. Note that the library-free form (EGFP only) is not amplified by the sequencing primers, thereby conserving NGS reads. (B) Volcano plot summarizing the MaPSy results. The x-axis represents the ratio of canonical splicing for reference versus variant (ref/variant ratio), whereas the y-axis indicates the maximal P-value across four experiments for each ref/variant oligo pair. Variants for which potential 3'ss (AG) have been added or deleted are marked accordingly. (C) Counts of high-confidence (HC) splicing-altering mutations, defined by significant splicing changes over four experiments with a greater than two-fold change. Categories include: "Unspliced" (significant increase in unspliced reads), "Noncanonical" (significant shift to noncanonical splice sites), and "Noncanonical+unspliced" (significant changes across both categories). This figure has been adapted with permission from Chiang et al.17. Please click here to view a larger version of this figure.

Supplementary File 1: Sequence of the pGint-CAMTA2 minigene. Please click here to download this File.

Supplementary File 2: Command lines for sequence and junction alignment. Please click here to download this File.

Discussion

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The intrinsic EGFP signal in the MaPSy construct enables fluorescence-based detection of exon skipping. If the sequence in the middle exon or introns promotes exon skipping, ligation of the first and third exons produces an EGFP signal detectable by FACS, making this a valuable method for identifying variants that influence exon skipping and facilitating microscopy-based visualization of splicing variants13. However, the amplicon-sequencing approach described herein does not capture exon-skipping products, as they lack library information and so they are not sequenced. Consequently, our method does not directly identify variants that induce exon skipping. Nevertheless, based on previous analyses, variants that reduce canonical splicing often trigger multiple noncanonical splicing events, including noncanonical splice site usage and exon skipping17,27. Therefore, our amplicon sequencing approach reliably identifies splicing variants even without detecting exon-skipping events.

Compared to FACS-based detection, amplicon sequencing is more robust in terms of identifying noncanonical splice site usage. FACS-based assays depend on middle exon inclusion to disrupt the EGFP reading frame, rendering them less sensitive to subtle alterations in splicing signals. Additionally, FACS-based experiments require stable integration of the MaPSy constructs into a fixed genomic site to maintain uniform expression levels across cells. Due to natural transcriptional variation, a secondary fluorescence signal is needed in FACS experiments to normalize EGFP expression. In contrast, MaPSy using amplicon sequencing allows for bulk preparation without the need for individual cell incorporation, as all constructs are driven by a CMV promoter and terminated by an SV40 polyadenylation signal. Although this linear PCR product lacks the complete modifications found in genomic DNA, all variants are tested under uniform conditions. This approach also minimizes the bias and mutations that can arise from bacterial amplification and genome integration, making it a faster and more efficient alternative.

MaPSy data on disease-relevant variants hold both clinical relevance and broader scientific potential, providing a basis for understanding the mechanisms of splice site selection. Machine learning and deep learning models have been employed to identify the regulatory rules governing variant-induced splicing defects, leveraging this data to predict novel splicing variants that impact human health4,11,12,13,14,27. Additionally, the MaPSy approach can be used to systematically vary sequence composition and splicing signal distances, offering a controlled method to explore how splicing is regulated14. Such approaches are particularly valuable for investigating splicing defects in cases where access to specific tissues is limited or when working with degraded RNA samples. Overall, MaPSy represents an invaluable approach for advancing our understanding of the splicing abnormalities that contribute to various diseases.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors declare no conflicts of interest.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Funding support for this work was provided by the Career Development Award, Multidisciplinary Health Cloud Research Program, Grand Challenge Seed Grant of Academia Sinica (AS-CDA-108-M03, AS-PH-109-01-3 and AS-GCS-113-L03), the Career Development Award of the National Health Research Institutes, Taiwan (NHRI-EX112-10908BC), and Excellent Young Scholar Research Grants and Ta-You Wu Memorial Award of National Science and Technology Council, Taiwan (MOST 112-2628-B-001-009-MY3 and 108-2118-M-001-013-MY5).

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
Direct-zol RNA MiniPrep Plus kitZymo ResearchR2072
Dulbecco’s Modified Eagle’s Medium (DMEM)Thermo Fisher Scientific11965084
Fetal Bovine Serum (FBS)Thermo Fisher Scientific26140079
L-GlutamineThermo Fisher ScientificA2916801 
Lipofectamine 3000 Thermo Fisher ScientificL3000015
Penicillin-StreptomycinThermo Fisher Scientific15140122
pGint plasmidAddgene24217
Phusion High-Fidelity DNA PolymeraseThermo Fisher ScientificF530L
QIAquick Gel Extraction KitQiagen28706
QIAquick PCR Purification Kit Qiagen28106
QIAxcel DNA Screening Kit (2400)Qiagen929004
SuperScript IV reverse transcriptaseThermo Fisher Scientific18090010

References

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,
  1. Alternative splicing as a regulator of development and tissue identity. Nat Rev Mol Cell Biol. 18 (7), 437-451 (2017).">Baralle, F. E., Giudice, J. Alternative splicing as a regulator of development and tissue identity. Nat Rev Mol Cell Biol. 18 (7), 437-451 (2017).
  2. A class of human exons with predicted distant branch points revealed by analysis of AG dinucleotide exclusion zones. Genome Biol. 7 (1), R1(2006).">Gooding, C., et al. A class of human exons with predicted distant branch points revealed by analysis of AG dinucleotide exclusion zones. Genome Biol. 7 (1), R1(2006).
  3. RNA splicing by the spliceosome. Annu Rev Biochem. 89, 359-388 (2020).">Wilkinson, M. E., Charenton, C., Nagai, K. RNA splicing by the spliceosome. Annu Rev Biochem. 89, 359-388 (2020).
  4. Predicting splicing from primary sequence with deep learning. Cell. 176 (3), 535-548.e24 (2019).">Jaganathan, K., et al. Predicting splicing from primary sequence with deep learning. Cell. 176 (3), 535-548.e24 (2019).
  5. Using positional distribution to identify splicing elements and predict pre-mRNA processing defects in human genes. Proc Natl Acad Sci U S A. 108 (27), 11093-11098 (2011).">Lim, K. H., Ferraris, L., Filloux, M. E., Raphael, B. J., Fairbrother, W. G. Using positional distribution to identify splicing elements and predict pre-mRNA processing defects in human genes. Proc Natl Acad Sci U S A. 108 (27), 11093-11098 (2011).
  6. Genomic basis for RNA alterations in cancer. Nature. 578 (7793), 129-136 (2020).">Calabrese, C., et al. Genomic basis for RNA alterations in cancer. Nature. 578 (7793), 129-136 (2020).
  7. Genome-wide discovery of human splicing branchpoints. Genome Res. 25 (2), 290-303 (2015).">Mercer, T. R., et al. Genome-wide discovery of human splicing branchpoints. Genome Res. 25 (2), 290-303 (2015).
  8. Large-scale analysis of branchpoint usage across species and cell lines. Genome Res. 27 (4), 639-649 (2017).">Taggart, A. J., et al. Large-scale analysis of branchpoint usage across species and cell lines. Genome Res. 27 (4), 639-649 (2017).
  9. Most human introns are recognized via multiple and tissue-specific branchpoints. Genes Dev. 32 (7-8), 577-591 (2018).">Pineda, J. M. B., Bradley, R. K. Most human introns are recognized via multiple and tissue-specific branchpoints. Genes Dev. 32 (7-8), 577-591 (2018).
  10. Profiling lariat intermediates reveals genetic determinants of early and late co-transcriptional splicing. Mol Cell. 82 (24), 4681-4699 (2022).">Zeng, Y., et al. Profiling lariat intermediates reveals genetic determinants of early and late co-transcriptional splicing. Mol Cell. 82 (24), 4681-4699 (2022).
  11. The human splicing code reveals new insights into the genetic determinants of disease. Science. 347 (6218), 1254806(2015).">Xiong, H. Y., et al. The human splicing code reveals new insights into the genetic determinants of disease. Science. 347 (6218), 1254806(2015).
  12. MMSplice: modular modeling improves the predictions of genetic variant effects on splicing. Genome Biol. 20 (1), 48(2019).">Cheng, J., et al. MMSplice: modular modeling improves the predictions of genetic variant effects on splicing. Genome Biol. 20 (1), 48(2019).
  13. A multiplexed assay for exon recognition reveals that an unappreciated fraction of rare genetic variants cause large-effect splicing disruptions. Mol Cell. 73 (1), 183-194.e8 (2019).">Chong, R., et al. A multiplexed assay for exon recognition reveals that an unappreciated fraction of rare genetic variants cause large-effect splicing disruptions. Mol Cell. 73 (1), 183-194.e8 (2019).
  14. Learning the sequence determinants of alternative splicing from millions of random sequences. Cell. 163 (3), 698-711 (2015).">Rosenberg, A. B., Patwardhan, R. P., Shendure, J., Seelig, G. Learning the sequence determinants of alternative splicing from millions of random sequences. Cell. 163 (3), 698-711 (2015).
  15. In silico tools for splicing defect prediction: a survey from the viewpoint of end users. Genet Med. 16 (7), 497-503 (2014).">Jian, X. Q., Boerwinkle, E., Liu, X. M. In silico tools for splicing defect prediction: a survey from the viewpoint of end users. Genet Med. 16 (7), 497-503 (2014).
  16. Benchmarking deep learning splice prediction tools using functional splice assays. Hum Mutat. 42 (7), 799-810 (2021).">Riepe, T. V., Khan, M., Roosing, S., Cremers, F. P. M., 't Hoen, P. A. C. Benchmarking deep learning splice prediction tools using functional splice assays. Hum Mutat. 42 (7), 799-810 (2021).
  17. Mechanism and modeling of human disease-associated near-exon intronic variants that perturb RNA splicing. Nat Struct Mol Biol. 29 (11), 1043-1055 (2022).">Chiang, H. L., et al. Mechanism and modeling of human disease-associated near-exon intronic variants that perturb RNA splicing. Nat Struct Mol Biol. 29 (11), 1043-1055 (2022).
  18. Degenerate minigene library analysis enables identification of altered branch point utilization by mutant splicing factor 3B1 (SF3B1). Nucleic Acids Res. 47 (2), 970-980 (2019).">Gupta, A. K., et al. Degenerate minigene library analysis enables identification of altered branch point utilization by mutant splicing factor 3B1 (SF3B1). Nucleic Acids Res. 47 (2), 970-980 (2019).
  19. ClinVar: public archive of interpretations of clinically relevant variants. Nucleic Acids Res. 44 (D1), D862-D868 (2016).">Landrum, M. J., et al. ClinVar: public archive of interpretations of clinically relevant variants. Nucleic Acids Res. 44 (D1), D862-D868 (2016).
  20. dbSNP-database for single nucleotide polymorphisms and other classes of minor genetic variation. Genome Res. 9 (8), 677-679 (1999).">Sherry, S. T., Ward, M. H., Sirotkin, K. dbSNP-database for single nucleotide polymorphisms and other classes of minor genetic variation. Genome Res. 9 (8), 677-679 (1999).
  21. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics. 26 (6), 841-842 (2010).">Quinlan, A. R., Hall, I. M. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics. 26 (6), 841-842 (2010).
  22. HISAT: a fast spliced aligner with low memory requirements. Nat Methods. 12 (4), 357-360 (2015).">Kim, D., Landmead, B., Salzberg, S. L. HISAT: a fast spliced aligner with low memory requirements. Nat Methods. 12 (4), 357-360 (2015).
  23. Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nat Biotechnol. 37 (8), 907-915 (2019).">Kim, D., Paggi, J. M., Park, C., Bennett, C., Salzberg, S. L. Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nat Biotechnol. 37 (8), 907-915 (2019).
  24. Twelve years of SAMtools and BCFtools. Gigascience. 10 (2), giab008(2021).">Danecek, P., et al. Twelve years of SAMtools and BCFtools. Gigascience. 10 (2), giab008(2021).
  25. NIH Image to ImageJ: 25 years of image analysis. Nat Methods. 9 (7), 671-675 (2012).">Schneider, C. A., Rasband, W. S., Eliceiri, K. W. NIH Image to ImageJ: 25 years of image analysis. Nat Methods. 9 (7), 671-675 (2012).
  26. Open Source Software in Life Science Research. , Woodhead Publishing. Cambridge. (2012).">Lind, R. Open Source Software for Image Processing and Analysis: Picture this with ImageJ. Open Source Software in Life Science Research. , Woodhead Publishing. Cambridge. (2012).
  27. SpliceAPP: an interactive web server to predict splicing errors arising from human mutations. BMC Genomics. 25 (1), 600(2024).">Huang, A. C., et al. SpliceAPP: an interactive web server to predict splicing errors arising from human mutations. BMC Genomics. 25 (1), 600(2024).

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

Splicing ErrorsIntronic VariantsRNA SplicingMassively Parallel AssaySplicing MinigenesExon JunctionSpliceosome AssemblyGenome EditingSplicing EfficiencyDisease Mutations

Related Articles