$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Functional genomics remains a crucial and challenging field, even in the post-genome era, with a significant portion of genomic regions still lacking functional characterization. A key reason for these challenges lies in the intricate nature of the human genome and transcriptome, which are characterized by pervasive transcription and universal alternative splicing in both protein-coding and non-coding regions1,2,3,4,5. This complexity has been extensively highlighted through RNA-targeted enrichment studies6,7,8,9. It is increasingly recognized that numerous novel genomic elements remain unmapped and functionally unannotated10,11,12in addition to the well-annotated segments of the human genome.
In recent years, advancements in reverse-genetics tools have been seen, such as RNA interference (RNAi) and CRISPR/Cas systems13,14,15,16,17,18,19,20,21. However, these methods have notable disadvantages, including non-specific effects and high costs associated with generating complex libraries of short hairpin RNAs (shRNAs) or single guide RNAs (sgRNAs)22,23,24,25. Additionally, these methods predominantly focus on annotated genomic elements. In contrast, forward genetics has the advantage of identifying novel, unannotated functional genomic elements. Given the complexity of the human genome and transcriptome, as well as the vast number of novel genomic elements, this capability of forward genetics is of significant importance.
Forward genetics cell libraries are frequently generated through random retroviral integration, making the efficient detection of genome-wide integration sites crucial for a forward genetics assay. Detecting these integration sites typically involves genome-walking tools26. While there have been significant advancements in these tools, few are optimized for next-generation sequencing (NGS) applications and subsequently applied in insertional mutagenesis studies. Insertional mutagenesis in cell lines often utilizes methods like inverse PCR (I-PCR) or linear amplification-mediated PCR (LAM-PCR) to detect integration sites27,28,29. These methods, however, usually require restriction enzyme digestion or ligation steps, which can reduce efficiency and introduce biases in the results.
To enhance the efficient detection of genome-wide integration sites in forward genetics screens, a method called the Insertion-based Screen for functional Elements and Transcripts (InSET) method was introduced here. By integrating genome-walking with NGS, InSET enables the high-throughput identification of lentivirus integration sites in lentivirus-based insertional mutagenesis libraries. Compared to existing methods27,28,29, InSET is relatively straightforward and simple, eliminating the restriction enzyme digestion or ligation steps, making it more convenient and accessible. Furthermore, although the existing forward genetics screen in mammalian cells is limited to very few haploid or near-haploid cell lines27,28,29,30,31, InSET has been proven to work in aneuploid cell lines32.
This study provides a detailed, step-by-step protocol for implementing the InSET method. Additionally, it demonstrates the application of InSET in insertional mutagenesis studies. InSET serves as a powerful tool for identifying novel exons within genomic elements, encompassing both protein-coding and non-coding RNAs in human cell lines.