$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
In cell biology, gene regulation plays a central role. At one or multiple steps along the gene expression pathway, genes have the potential to be regulated. These steps include transcription (initiation, elongation, and termination) as well as splicing, polyadenylation or 3’ end formation, RNA export, mRNA translation, and decay/localization of primary transcripts. At these steps, nucleic acid-binding proteins modulate gene regulation. Identification of binding sites for such proteins is an important aspect of studying gene control. Mutational analysis and phylogenetic sequence comparison have been used to discover regulatory sequences or protein-binding sites in nucleic acids, such as promoters, splice sites, polyadenylation elements, and translational signals1,2,3,4.
Pre-mRNA splicing is an integral step during gene expression and regulation. The majority of mammalian genes, including those in humans, have introns. A large fraction of these transcripts is alternatively spliced, producing multiple mRNA and protein isoforms from the same gene or primary transcript. These isoforms have cell-specific and developmental roles in cell biology. The 5’ splice site, the branch-point, and the polypyrimidine-tract/3’ splice site are critical splicing signals that are subject to regulation. In negative regulation, an otherwise strong splice site is repressed, whereas in positive regulation an otherwise weak splice site is activated. A combination of these events produces a plethora of functionally distinct isoforms. RNA-binding proteins play key roles in these alternative splicing events.
Numerous proteins are known whose binding site(s) or RNA targets remain to be identified5, 6. Linking regulatory proteins to their downstream biological targets or sequences is often a complex process. For such proteins, identification of their target RNA or binding site is an important step in defining their biological functions. Once a binding site is identified, it can be further characterized using standard molecular and biochemical analyses.
The approach described here has two advantages. First, it can identify a previously unknown binding site for a protein of interest. Second, an added advantage of this approach is that it simultaneously allows saturation mutagenesis, which would otherwise be labor intensive to obtain comparable information about sequence requirements within the binding site. Thus, it offers a quicker, easier, and less costly tool to identify protein binding sites in RNA. Originally, this approach (SELEX or Systematic Evolution of Ligands by EXponential enrichment) was used to characterize the binding site for the bacteriophage T4 DNA polymerase (gene 43 protein), which overlaps with the ribosome binding site in its own mRNA. The binding site contains an 8-base loop sequence, representing 65,536 randomized variants for analysis7. Second, the approach was also independently used to show that specific binding sites or aptamers for different dyes can be selected from a pool of approximately 1013 sequences8. In fact, this approach has been broadly used in many different contexts to identify aptamers (RNA or DNA sequences) for binding numerous ligands, such as proteins, small molecules, and cells, and for catalysis9. As an example, an aptamer can discriminate between two xanthine derivatives, caffeine and theophylline, which differ by the presence of one methyl group in caffeine10. We have extensively used this approach (SELEX or iterative selection-amplification) to study how RNA-binding proteins function in splicing or splicing regulation11, which will be the basis for the discussion below.
The random library: We used a random library of 31 nucleotides. The length consideration for the random library was loosely based on the idea that the general splicing factor U2AF65 binds to a sequence between the branch-point sequence and the 3’ splice site. On average, the spacing between these splicing signals in metazoans is in the range of 20 to 40 nucleotides. Another protein Sex-lethal was known to bind to a poorly characterized regulatory sequence near the 3’ splice site of its target pre-mRNA, transformer. Thus, we chose a random region of 31 nucleotides, flanked by primer binding sites with restriction enzyme sites to allow for PCR amplification and attachment of the T7 RNA polymerase promoter for in vitro transcription. The theoretical library size or complexity was 431 or approximately 1018. We used a small fraction of this library to prepare our random RNA pool (~1012-1015) for the experiments described below.