$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Here, we demonstrate a suite of error-corrected sequencing protocols that can be easily implemented to study mutations with low VAFs in different diseases. The most important factor is the incorporation of UMIs with each molecule before sequencing as they enable error-correction of the raw reads. The methods described here allow researchers to incorporate customized UMIs to both commercially available gene panels and self-designed gene-specific oligos.
Standard NGS protocol precludes the detection of mutations with VAF below 2% due to the sequencing error rate, and this limits the application of NGS in studies where the detection of rare variants is crucial. By circumventing the standard NGS error rate, ECS enables sensitive detection of these raw variants. For instance, detection of pathogenic mutations when these mutations first arise (therefore having low VAF) is imperative to inform early intervention of the disease14,15. In leukemia research, the detection of minimal residual disease (residual leukemic cells post-treatment) informs risk stratification and could be used to inform treatment options in a manner that binary flow cytometric assessments cannot. In addition, ECS is applicable to detect circulating tumor nucleic acid and to evaluate metastatic potential in solid tumor patients by assessing for the presence/absence as well as the variant burden of certain mutations that are characteristics of the primary tumor16.
As demonstrated in Table 1, the power of using binomial distribution-based position-specific error model to call variants depends largely on the number of sequenced libraries as well as the depth of sequencing used to build the error model. The robustness of the error model increases with higher number of samples and more sequencing depth. It is recommended to use at least 10 sequenced samples with an average of error-corrected read coverage of 3000x per sample to build an error profile for each sample. The position-specific approach is similar to MAGERI, but instead of using an aggregate error rate for all six different substitution types (A>C/T>G, A>G/T>C, A>T/T>A, C>A/G>T, C>G/G>C, C>T/G>A)13, we model each substitution independently at every position. For instance, an error rate of C>T at a given genomic position is different from another position. Our approach also takes into account a sequencing batch effect, as the base substitution rate observed in one sequencing run might be different from another run. Hence it is important to model each position for all substitution types especially when samples from different sequencing runs are pooled to build the model.
An important consideration when designing an ECS experiment is the desired detection threshold. The beauty of NGS studies is that they can be easily scaled in terms of genes/targets of interest, detection threshold (dictated by depth of sequencing), and number of individuals queried. For example, if the researchers are interested to find rare mutations in two amplicons with a detection threshold of 0.0001, they can pool maximally 75 samples in a single sequencing run using MiSeq V2 chemistry which outputs up to 15 million reads (2 amplicons * 10,000 molecules * 10 reads for error-correction * 75 samples = 15 million sequencing reads). Researchers can vary the number of molecules going into sequencing or the number of pooled samples in a single sequencing run to adjust the detection threshold. In our studies, we aimed to find mutations with a detection threshold of 0.0001 VAF (1:10,000) using the Illumina gene panel. We routinely use 250 ng of starting DNA to ensure that sufficient molecules are captured in order to achieve the aforementioned detection threshold. Researchers can opt to start with lower amount of DNA (50 ng is recommended) if the desired detection limit is >0.001 VAF.
As the UMIs are appended onto the i5 indexes, sequencing settings have to be amended accordingly. For example, we used 16 N UMIs, and the sequencing settings were 2x144 paired end reads, 8 cycles of Index 1 and 16 cycles of index 2 as opposed to the usual 8 cycles of Index 2. The increase in Index 2 cycle is compensated by a decrease in the total number of cycles allocated to the reads. If researchers opt to use 12N UMIs10,17, the settings should be changed to 12 cycles of Index 2.
This UMI-based sequencing method is optimized to correct for sequencing errors. It remains suboptimal in dealing with PCR jackpotting, which is an issue for all amplification-based method. We performed rounds of post-sequencing and post-bioinformatics validation using ddPCR, and we hardly detect any false positives due to PCR jackpotting. Nonetheless, it is recommended that researchers conduct the experiments using high fidelity polymerase to ensure low amplification errors.