$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Mutagenesis has long been employed in the laboratory to study the properties of biological systems and their evolution, and to produce mutant proteins or organisms with enhanced or novel functions. While early approaches relied on methods which produce random mutations in organisms, the advent of recombinant DNA technology enabled researchers to introduce select changes to DNA in a site-specific manner, i.e., site-directed mutagenesis1,2. With current techniques, typically using mutagenic oligonucleotides in a polymerase chain reaction (PCR), it is relatively facile to create and assess small numbers of mutations (e.g., point mutations) in a given gene3,4. It is far more difficult however when the goal approaches, for example, the creation and assessment of all possible single-site (or higher-order) mutations.
While much has been learned from early studies attempting to assess large numbers of mutations in genes, the techniques used were often laborious, for example requiring the assessment of each mutation independently using nonsense suppressor strains5-7, or were limited in their quantitative ability due to the low sequencing depth of Sanger sequencing8. The techniques used in these studies have largely been supplanted by methods utilizing high-throughput sequencing technology9-12. These conceptually simple approaches entail creating a library comprising a large number of mutations, subjecting the library to a screen or selection for function, and then deep-sequencing (i.e., on the order of >106 sequencing reads) the library obtained before and after selection. In this way, the phenotypic or fitness effects of a large number of mutations, represented as the change in population frequency of each mutant, can be assessed simultaneously and more quantitatively.
We previously introduced a simple approach for assessing libraries of all possible single amino acid mutations in proteins (i.e., whole-protein saturation mutagenesis libraries), applicable to genes with a length longer than the sequencing read length11,13: First, each amino acid position is randomized by site-directed mutagenic PCR. During this process, the gene is split into groups composed of contiguous positions with a total length accommodated by the sequencing platform. The mutagenic PCR products for each group are then combined, and each group independently subjected to selection and high-throughput sequencing. By maintaining a correspondence between the location of mutations in the sequence and the sequencing read length, this approach has the advantage of maximizing sequencing depth: while one could simply sequence such libraries in short windows without splitting into groups (e.g., by a standard shotgun sequencing approach), most reads obtained would be wild-type and thus the majority of sequencing throughput wasted (e.g., for a whole-protein saturation mutagenesis library of a 500 amino acid protein sequenced in 100 amino acid (300 bp) windows, at minimum 80% of reads will be the wild-type sequence).
Here, a protocol is presented which utilizes high-throughput sequencing for the functional assessment of whole-protein saturation mutagenesis libraries, using the above approach (outlined in Figure 1). Importantly, we introduce the usage of orthogonal primers in the library cloning process to barcode each sequence group, which allows them to be multiplexed into one library, subjected simultaneously to screening or selection, and then de-multiplexed for deep sequencing. Since the sequence groups are not subjected to selection independently, this reduces the workload and ensures that each mutation experiences the same level of selection. TEM-1 β-lactamase, an enzyme which confers high-level resistance to β-lactam antibiotics (e.g., ampicillin) in bacteria is used as a model system14-16. A protocol is described for the assessment of a whole-protein saturation mutagenesis library of TEM-1 in E. coli under selection at an approximate serum level for a clinical dose of ampicillin (50 µg/ml)17,18.