$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
CORALINA can be used to generate large scale gRNA libraries by controlled nuclease digestion of target DNA and bulk cloning of resulting double stranded fragments. Statistical inference indicates that many more than 107 individual gRNA sequences have already been successfully cloned using the protocol at hand9. CORALINA can be customized in multiple ways. The choice of template DNA defines the target region and the maximal complexity of the generated library. Using this protocol, CORALINA libraries have previously been generated from human and mouse genomic DNA9. Representative results presented here depict the generation of a CORALINA library from purified BAC DNA. Further customization can be achieved by the choice of gRNA expression vector and linker sequences. We have previously tested three different pairs of linker lengths for Gibson assembly with little variations in efficiency9.
Due to their origin from bulk digested DNA, protospacer of CORALINA gRNAs are usually not exactly 20 bp in length, but show a length distribution with a mean that depends both on the parameters of the MNase digestion as well as the size of the excision made from the PAGE gels. The representative example shown in Figure 2B and C, depicts fragments with a median length between 19 and 27 bp. In our experience, the length of the fragments is faithfully preserved by the generated gRNA protospacer9. While fragments shorter than 20 bp should be avoided due to higher off-target rate of resulting gRNAs, longer fragments are likely much less of a problem for downstream applications, since it has been demonstrated that gRNAs with protospacers as long as 45 bp are still functional9.
The two most critical steps in the CORALINA protocol are the size selection of MNase-digested fragments and cloning steps. Generation of fragments that are too short (e.g. average below 18 bp) or incorporation of too many empty gRNA expression vectors will render the library useless. Thus, it is important to optimize the MNase digestion step (Figure 2A), to monitor excision (Figure 2B, C), check for complete digestion of the gRNA vector backbone and including no fragment controls throughout the protocol. Special care has also to be taken to preserve the representation of the gRNA library. One common bottleneck of library generation in general is the efficient transfer of plasmids into bacteria for amplification. Thus, large quantities of bacteria with excellent competency and a large number of individual electroporation events are necessary to achieve a high number of gRNA clones.
New strategies for gRNA library production will be necessary to harvest the full potential of CRISPR-based screening approaches over the next decades. There is a significant demand for cost-effective, simple and customizable methods to generate large-scale libraries, a pre-requisite to make screening amenable to a larger number of model systems and different CRISPR-based engineering approaches. CORALINA is providing a first step toward this. The potential uses are manifold, especially to produce comprehensive libraries of genomes, cDNA derived libraries of less common model systems, highly focused libraries and experimental set-ups in which different CRISPR proteins (with differing PAM requirements) are used in combination.
Unlike other methods, CORALINA generates all possible gRNAs from the input DNA. However, one drawback of the method is that gRNAs lacking the required PAM sequence are also included in the library, a feature that it shares with a second enzymatic method for gRNA library generation, CRISPR-EATING (Table 1). The choice of the ideal method for gRNA library generation depends on the specifications of the planned screening experiment, especially the nature (genic, regulatory, intergenic) and size of the target region (single locus, multiple regions, genome-wide). We see a special upside in using CORALINA when a large number of non-coding or regulatory regions are to be analyzed, if there is incomplete or unreliable sequence information (exotic model systems, mixtures of species (e.g. microbiomes) or experimentally obtained input), if different CRISPR endonucleases are combined or if saturating analysis is performed on a short and defined locus (e.g. represented by BACs).