$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
1. Generate Virus Stocks
Note: A flow chart of the wet bench aspect of this protocol is depicted in Figure 1. The details of viral stock production and subsequent infection of tissue culture cells will generally apply to different types of retroviruses. For some experiments, the target cell may not express the endogenous viral receptor(s), and in such cases the construction of pseudotyped retroviral particles harboring heterologous viral envelope glycoprotein, e.g. the G glycoprotein from vesicular stomatitis virus (VSV-G), will be required for infection 44,45.
Note: Precaution should be taken when working with HIV-1. Though specific guidelines will vary from institution to institution, all virus-based work should be conducted in a dedicated, operator restricted biological safety cabinet (typically referred to as a tissue culture hood). Proper personal protective equipment that includes face protection, shoe covers, a double glove layer, and a full-body coverall suit should be worn at all times. All liquid waste resulting from virus-related experiments should be inactivated with bleach (10% final concentration), and all waste including solids should be autoclaved prior to disposal.
- One day prior to transfection, plate 3.3 x 106 HEK293T cells in 10 ml of Dulbecco's Modified Eagle Medium (DMEM) supplemented with 10% (v/v) fetal bovine serum and 1% (v/v) penicillin/streptomycin (10,000 U/ml stock) in each of five 100 mm dishes.
Note: Supplemented-DMEM is referred to as DMEM-FPS from this point on.
- On the subsequent day, transfect the cells with 10 µg of plasmid carrying full-length retroviral molecular clones or 9 µg of envelope-deleted single-round vectors with 1 µg of a VSV-G expression construct using commercially available transfection reagents or calcium phosphate.
- Incubate the cells at 37 °C in a humidified cell culture incubator with 5% CO2 (this condition hereafter referred to as "tissue culture incubator"). After approximately 48 hr, harvest the virus-containing cell media using a volumetric pipette and pass it through a 0.45 µm filter by gravity flow.
- Concentrate the virus by ultracentrifugation at 200,000 x g for 1 hr at 4 °C. Resuspend the virus pellet in 500 µl DMEM-FPS containing 20 U DNase, and incubate for 1 hr at 37 °C.
Note: The DNase step helps to reduce the recovery of unwanted plasmid sequences by eliminating the brunt of plasmid DNA that persists from the transfection procedure.
- Determine p24 concentration 46 using an HIV-1 p24 antigen capture kit as per manufacturer's instructions.
Note: Virus concentration can also be determined by reverse transcriptase activity assay 47,48. Alternatively, the level of functional virus can be determined by measuring MOI. This is most readily done using fluorescence-activated cell sorting with viruses that express fluorescent reporter genes such as enhanced green fluorescent protein. MOI determination may be particularly useful when working with primary cells that may not support the same level of infection as optimized cell lines.
2. Infect Cells with Virus
- Plate 3.0 x 105 HEK293T cells per well in a 6-well plate in 2.5 ml DMEM-FPS and incubate overnight in a tissue culture incubator.
Note: The number of unique integration sites recovered with this protocol is directly proportional to the number of cells and quantity of active virus used in the infection.
- Infect cells with a final viral p24 concentration of 500 ng/ml in a final volume of 500 µl fresh DMEM-FPS for 2 hr in a tissue culture incubator, then add 2 ml DMEM-FPS pre-warmed to 37 °C per well and continue incubation.
- At 48 hr post-infection, remove the media and wash the cells with 2 ml phosphate-buffered saline (PBS). Add 0.5 ml trypsin-EDTA pre-warmed to 37 °C, and after a few sec visually inspect the wells for cell dislodgement.
- Add 2 ml pre-warmed DMEM-FPS and resuspend the cells by gentle up/down pipetting with a volumetric pipette ~10 times. Transfer the solution to a 75 cm2 tissue culture flask containing 18 ml pre-warmed DMEM-FPS, and incubate the cells in a tissue culture incubator.
- After minimally five days from the start of the infection, collect the cells by removing the media, wash with 5 ml PBS, add 2 ml pre-warmed trypsin-EDTA, and resuspend with 5 ml pre-warmed DMEM-FPS by pipetting. Centrifuge the solution for 5 min at room temperature at 2,500 x g, and discard the supernatant.
Note: Although integration under these conditions plateaus at about 48 hr post-infection 49,50, the additional 3 days of culture are required to sufficiently dilute the concentration of unintegrated DNA molecules that result from cell-based DNA recombination or viral-mediated autointegration.
- Extract genomic DNA from the cell pellet using a commercially available kit (e.g., see 51). Elute the DNA from the supplied ion exchange column with 200 µl of 10 mM Tris-HCl, pH 8.5.
Note: An aliquot of cells should be apportioned at 48 hr post-infection (Step 2.3) for an infectivity assay to ensure proper virus infection prior to NGS.
3. Fragment Genomic DNA by Sonication or by Restriction Enzyme Digest
Note: Sonication fragments genomic DNA in a virtually sequence-independent manner and is thus the preferable mode of fragmentation when sequencing samples with a low expected recovery rate (e.g., infected patient cells or infections initiated at relatively low MOI). Furthermore, sonication allows one to distinguish PCR duplicates of a particular integration site sequence from unique integrations at the same site, which is critical to distinguish the clonal expansion of provirus-containing cells in infected patients (see Step 11 below) 39,52-54.
Note: The DNA should be cleaved immediately downstream from the upstream LTR to diminish amplification of internal viral sequences during LM-PCR. The restriction enzyme BglII that lies 43 bp downstream from the upstream U5 sequence and that is incompatible for subsequent ligation with MseI-generated DNA ends works well with many HIV-1 strains (Figure 1B). When preparing DNA by sonication, the internal-cleaving restriction enzyme should be applied after linker ligation (see Figure 1C-E and Step 4.3 below).
- For sonication, mix 10 µg of genomic DNA in nuclease-free water to a final volume of 120 µl. Sonicate using parameters for an average break size of 500 bp (two rounds of the following parameters: duty cycle: 5%; intensity: 3; cycles per burst: 200; time: 80 sec).
- Purify sonicated DNA using a PCR purification kit. Repair the DNA ends using a DNA end-repair kit and purify the DNA using a PCR purification kit. A-tail the DNA using Klenow exo- enzyme and purify the A-tailed DNA using a PCR purification kit. Refer to 51,52 for additional details of kit usage.
- For restriction endonuclease digestion, cut 10 µg of genomic DNA overnight at 37 °C in a volume of 100 µl with buffer supplied by the manufacturer and a cocktail of enzymes (100 U each) that generate 5'-TA overhangs, as well as an incompatible enzyme such as BglII that cleaves downstream from the upstream viral LTR. Purify the DNA the next day using a PCR purification kit.
Note: None of the restriction enzymes should cut within the terminal ~30 bp of the viral DNA end that is amplified by the LM-PCR protocol. This protocol specifically amplifies the U5 end of HIV-1 DNA.
4. Anneal Linker Oligonucleotides and Ligate to Fragmented Genomic DNA
Note: Prepare an asymmetric linker containing an overhang that is compatible with the above DNA fragments (see Table 1 for the sequences of oligonucleotides utilized in this protocol). The linker to be used with sonicated DNA must contain a compatible T-3' overhang, while the linker for MseI-digested DNA must contain a compatible 5'-TA overhang (Figure 1). The short linker strand must additionally contain a non-extendable chemical modification, such as 3'-amine, to constrain the subsequent amplification reactions toward the DNA of interest.
Note: When preparing multiple different integration site libraries in parallel and/or when multiplexing unique samples on the same sequencing run, it is recommended to use unique linkers for each sample to limit the potential for sample cross-contamination during PCR. This additionally implies the use of unique linker primers for each sample during semi-nested PCR (described below). Unique linker strands and linker primers may be designed by scrambling the linker oligonucleotide sequences listed in Table 1 while maintaining similar overall %GC content and applicable overhang positions.
- Anneal the short and long linker strands in 35 µl of 10 mM Tris-HCl, pH 8.0-0.1 mM EDTA (final concentration of 10 µM of each oligonucleotide) by heating to 90 °C and slowly cooling to room temperature in steps of 1 °C per min.
- Prepare at least four parallel ligation reactions per genomic DNA sample, which contain 1.5 µM ligated linker, 1 µg fragmented DNA, and 800 U T4 DNA ligase in 50 µl. Ligate overnight at 12 °C. Purify the next day with a PCR purification kit.
- For samples prepared by sonication, digest the purified ligation reaction with 100 U of a restriction enzyme that cleaves downstream from the upstream LTR (e.g., BglII for HIV-1) under the manufacturer's recommended conditions overnight. Purify the DNA using a PCR purification kit.
5. Amplify Viral LTR-Host Genomic DNA Junctions by Semi-nested PCR
Note: To ensure for optimal library diversity, at least 4-8 parallel PCRs, depending on the DNA concentration of the recovered ligation reaction, should be prepared for each sample for both PCR rounds. DNA template concentration should be quantified by spectrophotometry. In this protocol the first and second rounds of PCR employ nested LTR-specific primers, but the same linker-specific primer is used for both rounds (Table 1). The second round LTR-specific primer and the linker-specific primer encode adapter sequences for DNA clustering as well as sequencing primer-binding sites. The nested LTR-specific primer also encodes a 6 nt index sequence, which can be varied among different primers for multiplexing libraries within the same sequencing run.
- Prepare first round PCRs containing the ingredients per tube as listed in Table 2.
Note: The linker-specific primer harbors 22 nt of complementarity to the linker, a melting temperature of 53 °C, a GC-content of 45%, and its 3' end is located 15-16 bp upstream from the 3' termini of the different linker long strands (Table 1). The first round 27 nt LTR primer has a melting temperature of 59 °C, a GC-content of 48%, and its 3' end is located 34 bp upstream from the HIV-1 U5 terminus. The region of the second round 26 nt LTR primer that is complementary to the HIV-1 LTR has a melting temperature of 60 °C, a GC-content of 50%, and its 3' end is located 18 bp upstream from the viral U5 terminus. It is recommended that oligonucleotide melting temperature and GC-content should mimic these parameters if users design PCR primers with altered sequences (including for use with other retroviruses) 21.
- Run first PCR round under the following thermocycler parameters: One cycle: 94 °C for 2 min; 30 cycles: 94 °C for 15 sec, 55 °C for 30 sec, 68 °C for 45 sec; one cycle: 68 °C for 10 min.
- Pool reactions and purify using a PCR purification kit. Prepare second round PCRs containing the ingredients per tube as per Table 3. Run the second round of PCR using the thermocycler parameters described in Step 5.2. Pool the reactions and purify the DNA using a commercial PCR purification kit following the manufacturer's instructions.
Note: A variety of recommended index sequences compatible with DNA clustering NGS are available 71.
6. Perform QC and NGS (Typically Completed by a Sequencing Facility)
- (QC assay #1) Confirm Step 5.3 library DNA concentration using a fluorometer 55. Briefly, prepare standards and experimental samples in a final volume of 200 µl nuclease-free water. Vortex tubes for 2-3 sec, incubate at room temperature for 2 min, and then read the samples in the fluorometer.
Note: Samples should contain a minimum concentration of 2 nM library DNA in a minimum volume of 15 µl.
- (QC assay #2) Confirm DNA fragment size distribution using a tape-based assay 56.
Note: An ideal distribution is a relatively broad DNA peak centering around 500 bp in length. If a significant amount of material is larger than 1 kb, then it is recommended to incorporate a size-selection procedure to eliminate longer DNA species, which will impede bridge amplification during clustering. By contrast, if a significant peak is apparent around 100 to 200 bp, a primer dimer may have formed during PCR. In this case the procedure should be optimized to minimize the formation of primer dimers.
- (QC assay #3) Confirm proper incorporation of adapters into DNA library by quantitative PCR 57.
- Perform NGS following the manufacturer's application literature. Utilize a spike-in of 10% (w/w) ΦX174 DNA, which will optimize real time quality metrics by providing balanced base composition to the sequencing run.
Note: Integration site sequencing experiments are typically subjected to single end 150 bp (SE150) or paired-end 150 bp (PE150) sequencing. PE150 is particularly useful to capture the linker attachment point on each DNA molecule (e.g., when scrutinizing integration sites for evidence of host cell clonal expansion).
7. Use a Customized Python or PERL Script to Parse Sequencing Data for LTR-containing Sequences, Crop away LTR and Linker Sequences, and Map to Reference Genome with BLAT
- Scan FASTA files for LTR-containing sequence reads, crop LTR and linker sequences away from host genomic DNA sequence, and export these sequences to a new FASTA file. Map cropped reads to both a reference genome (e.g. human genome versions hg19 or GRCh38) and the viral genome using BLAT 58, with output integration site coordinates exported to a separate .txt file, using the following settings:
stepSize = 6, minIdentity = 97, and maxIntron = 0
- Parse the BLAT output .txt file, remove autointegrations (i.e. evidence that the LTR end has integrated into an internal region of the viral DNA genome) and other sequences mapping to the HIV-1 genome, and create a separate output .txt file in which all duplicate integration sites have been condensed into single, unique coordinate hits.
8. Create .bed Files Containing 15-Nt Intervals Surrounding Integrations, Convert These to FASTA Files, and Construct Sequence Logos to Display Base Preferences Surrounding Integration Sites
- Create .bed files that list an interval of bases for each integration site. At least 15 bases (5 upstream and 10 downstream) are suggested for sequence logo generation. Generate a FASTA file from these .bed files by using the fastaFromBed function from BEDTools 59 and this command:
fastaFromBed -fi /directory/to/reference/genome/ -name -s -bed 15_base_pair_file.bed -fo output_file.fasta
Note: The invariant viral 5'-CA-3' dinucleotide is joined to host DNA during integration, and verifying the junction of the LTR terminus to cellular DNA is an important initial filter to identify bona fide integration sites. We additionally compile sequence logos from this host DNA sequence population to verify the experimental results. As retroviruses display signature base preferences surrounding their integration sites 14,15, the sequence logos serve to validate that the mapped genomic sites arose through IN-mediated integration as compared to other recombination mechanisms such as non-homologous DNA end joining 60,61.
- Use WebLogo 3 (http://weblogo.threeplusone.com/create.cgi) to create sequence logos from the FASTA files. Click 'Choose File' to upload FASTA file, and use the following settings: Output format, PDF (vector); Logo size, large; First position number, -5; Logo range, -5 to 5; Y-axis scale, 0.1, Y-axis tic spacing, 0.5, Color scheme, classic (NA).
9. Create Central Base Pair .bed Files, Check for Sample Cross-Contamination, and Map the Distribution of Unique Integration Sites Relative to Pertinent Genomic Features
- Since retroviral integration occurs in a staggered fashion across the tDNA strands, adjust the precise coordinates of integration sites to reflect the central bp of the target site duplication for correct mapping of genomic distribution relative to genomic features.
- Therefore, for 5 bp duplicating viruses like HIV-1, create a .bed file with the central bp offset from the integration site by two bases downstream for integrations mapping to the plus strand, and two bases upstream for integrations mapping to the minus strand.
- To check for sample cross-contamination, calculate the number of integration sites common among the different libraries by using the BEDTools intersect function to intersect central bp .bed files for two different samples and by following this command:
bedtools intersect -a central_basepair_1.bed -b central_basepair_2.bed -f 1.00 -r -s > overlap1v2.txt
- Count the number of lines within the output overlap1v2.txt file in order to quantify the exact number of sites common among the two libraries by using the following command:
wc -l overlap1v2.txt
- Download the RefSeq annotation .bed file for the version of reference genome that was used for integration site mapping from the UCSC Genome Annotation Database (e.g. http://hgdownload.cse.ucsc.edu/goldenPath/hg38/database) 62.
- Calculate the number of integration sites falling within RefSeq genes by using the BEDTools intersect function to intersect the central base pair .bed file that was generated for the sample with the RefSeq .bed file following this command:
bedtools intersect -a central_basepair_1.bed -b RefSeq_hg38.bed -u > RefSeq_sample1.bed
- Count the number of lines within the output RefSeq_sample1.bed file in order to quantify the exact number of sites falling in RefSeq genes by using the following command:
wc -l RefSeq_sample1.bed
- Repeat steps 9.3 and 9.4 for mapping integration sites to any other annotation of interest for which an interval .bed file is available. Download the most current CpG island annotation .bed file for the reference genome of interest from the UCSC Genome Annotation Database as directed in Step 9.4.
- Calculate the number of integration sites falling within a certain distance (illustrated in this example is a 5 kb window) of CpG islands by using the BEDTools window function and following this command:
bedtools window –w 2500 central_basepair_1.bed -b CpG_hg38.bed -u > CpG_sample1.bed
- Count the number of lines within the output CpG_sample1.bed file in order to quantify the exact number of sites falling within 2.5 kb upstream or downstream of CpG islands by using the following command:
wc -l CpG_sample1.bed
- Repeat steps 9.6 and 9.7 for mapping integration sites nearby TSSs. Generate an alternate version of the RefSeq.bed file, where genomic coordinates mapping to more than one gene have been adjusted to reflect only a single gene present at that position. This prevents overestimation of gene density surrounding integration sites. Calculate the gene density in the 1 Mb region surrounding each integration site by using the BEDTools window function and following this command:
bedtools window -w 500000 central_basepair_1.bed -b RefSeq_hg38_NonRedundant.bed -u > GeneDensity_sample1.bed
- Calculate the average gene density for all integrations in the dataset by following this command:
awk '(sum+=$7) END ( print "Average = ",sum/NR)' GeneDensity_sample1.bed
10. Statistically Compare Integration Site Distributions among Samples Using Two-tailed Fisher's Exact Test and Two-tailed Wilcoxon Rank Sum Test in R
Note: Use Fisher's exact test for comparing the proportion of integration sites within RefSeq genes or within a window of CpG islands or TSSs, but use the Wilcoxon rank sum test for comparing the distribution in gene density surrounding the integration sites. The R program is available at http://www.r-project.org/.
Two-tailed Fisher's exact test:
- Using the numbers calculated as instructed in Steps 9.4 and 9.7, create matrices for each comparison in R of observed occurrences (integrations within an annotation or within a window surrounding an annotation) versus remaining sites by following this command:
(annotation_of_interest <- matrix(c(SampleA#in, SampleA#remaining, SampleB#in, SampleB#remaining),nrow=2,dimnames=list(c('Center', 'Remainder'),c('SampleA', 'SampleB'))))
- Calculate the P value for the comparison by two-tailed Fisher's exact test with the following command:
fisher.test(annotation_of_interest, alternative = 'two.sided')$p.value
Two-tailed Wilcoxon rank sum test:
- Create a tab-delimited .txt file in which each column contains the sample name in the top cell, followed below by the gene density values for all integration sites in that library (obtained from the .bed file generated in Step 9.9). Import this tab-delimited .txt file into R using the following command and navigating to the correct file directory:
FILENAME <- as.data.frame(read.delim(file.choose(), header=T,check.names=FALSE, fill=TRUE,sep='\t'))
- Calculate the P value for the comparison by two-tailed Wilcoxon rank sum test with the following command:
wilcox.test(FILENAME$SampleA,FILENAME$SampleB,alternative ='two.sided', paired = F, exact = T)$p.value
Note: P values can be calculated only down to a certain (extremely low) limit in R, after which zero will be returned by the program. For massively different samples that yield a P = 0 in R, estimate the P value as <2.2 x 10-308.
11. Examine Raw Sequencing Data for Evidence of Clonal Expansion of Cells Containing Integrated Viral DNA
Note: A small potential exists for more than one integration at the exact same nt in the reference genome. Alternatively, a single integration event may become redundantly present in sequencing data due to the use of PCR during library preparation and/or by cell duplication prior to DNA preparation. Recent analyses of genomic DNA from HIV-infected patients have distinguished these possibilities by identifying unique sonication shear points/linker attachment points (which can only arise prior to PCR) within DNA sequences containing identical integration sites 52-54. There is currently a debate as to whether proviruses harbored within clonally expanded cells contribute to the latent viral reservoir, and thus it is of particular interest to characterize their level of expansion when studying integration sites in human patients.
- Similar to the procedure listed in Step 8.1, generate .bed files listing an interval of bases extending, in this case, 25 nt downstream from each unique integration site (upstream bases are unnecessary here). Generate a FASTA file from these .bed files (as instructed in Step 8.1) by using the fastaFromBed function from BEDTools and following this command:
fastaFromBed -fi /directory/to/reference/genome/ -name -s -bed 25_base_pair_file.bed -fo output_file.fasta
Note: To improve the specificity of each search it is recommended to extract at least 25 nt downstream from each integration site for clonal expansion analyses.
- Preferably using a customized script, search the raw sequence data FASTA file for all strings containing an exact match to the 25 nt downstream from each unique integration site, and deposit these sequences into a new file. Trim LTR and linker sequences from the raw strings. Merge PE sequence reads by converting reads to the reverse complement, trimming LTR and linker sequences, and then assigning read2 strings to their read1 pair if the strings share at least 20 overlapping nt.
- Scan the linker attachment points of each integration site block. Classify each integration as "clonally expanded" if linker attachment points are ≥3 bp apart.
Note: A protocol for clonal expansion analysis without merging sequence reads has been described 52.
Note: Fragmentation of the genome at the exact same location by sonication leads to an underestimation of the extent of clonal expansion, and methods to correct the resulting experimental bias have been described 63,64.