$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Most human genes undergo Alternative Splicing (AS) to produce multiple isoforms with distinct activities, which has greatly increased the coding complexity of the genome1,2. AS provides a major mechanism to regulate gene function, and it is tightly regulated through diverse pathways in different cellular and developmental stages3,4. Because splicing misregulation is a common cause of human disease5,6,7,8, targeting splicing regulation is becoming an attractive therapeutic route.
According to a simplified model of splicing regulation, AS is mainly controlled by Splicing Regulatory cis-Elements (SREs) in pre-mRNA that function as splicing enhancers or silencers of alternative exons. These SREs specifically recruit various trans-acting protein factors (i.e. splicing factors) that promote or suppress the splicing reaction3,9. Most trans-acting splicing factors have separate sequence-specific RNA binding domains to recognize their targets and effector domains to control splicing. The best-known examples are the members of the serine/arginine-rich (SR) protein family that contain N-terminal RNA Recognition Motifs (RRMs), which bind exonic splicing enhancers, and C-terminal RS domains, which promote exon inclusion10. Conversely, hnRNP A1 binds to exonic splicing silencers through the RRM domains and inhibits exon inclusion through a C-terminal glycine-rich domain11. Using such modular configurations, researchers should be able to engineer artificial splicing factors by combining a specific RNA-Binding Domain (RBD) with different effector domains that activate or inhibit splicing.
The key of such a design is to use an RBD that recognizes given targets with programmable RNA binding specificity, which is analogous to the DNA-binding mode of the TALE domain. However, most native splicing factors contain RRM or K Homology (KH) domains, which recognize short RNA elements with weak affinity and thus lack a predictive RNA-protein recognition "code"12. The RBD of PUF repeat proteins (i.e. the PUF domain) has a unique RNA recognition mode, allowing for the redesign of PUF domains to specifically recognize different RNA targets13,14. The canonical PUF domain contains eight repeats of three α-helices, each recognizing a single base in an 8-nt RNA target. The side chains of amino acids at certain positions of the second α-helix form specific hydrogen bonds with the Watson-Crick edge of the RNA base, which determines the RNA binding specificity of each repeat (Figure 1A). The code for RNA base recognition of the PUF repeat is surprisingly simple (Figure 1A), allowing for the generation of PUF domains that recognize any possible 8-base combination (reviewed by Wei and Wang15).
This modular design principle allows for the generation of an Engineered Splicing Factor (ESF) that consists of a customized PUF domain and a splicing modulation domain (i.e. an SR domain or a Gly-rich domain). These ESFs can function as either splicing activators or as inhibitors to control various types of splicing events, and they have proven useful as tools to manipulate the splicing of endogenous genes related to human disease16,17. As an example, we have constructed PUF-Gly-type ESFs to specifically alter the splicing of the Bcl-x gene, converting the anti-apoptotic long isoform (Bcl-xL) to the pro-apoptotic short isoform (Bcl-xS). Shifting the ratio of the Bcl-x isoform was sufficient to sensitize several cancer cells to multiple anti-cancer chemotherapy drugs16, suggesting that these artificial factors may be useful as potential therapeutic reagents.
In addition to controlling splicing with known splicing effector domains (e.g., an RS or Gly-rich domain), the engineered PUF factors can also be used to examine the activities of new splicing factors. For example, using this approach, we have demonstrated that the C-terminal domain of several SR proteins can activate or inhibit splicing when binding to different pre-mRNA regions18, that the alanine-rich motif of RBM4 can inhibit splicing19, and that the proline-rich motif of DAZAP1 can enhance splicing20,21. These new functional domains can be used to construct additional types of artificial factors to fine-tune splicing.