$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Analyzing the whole process of gene regulation is necessary to fully understand cellular adaptations in response to a specific stimulus or treatment. Combining total RNA-seq, 4sU-seq, ribosome profiling, and ChIP-seq at different time points leads to a comprehensive analysis of the main processes of gene regulation over time. A profound understanding of biological processes is required to define the experimental setup as well as optimal time points.
Since methods to study gene regulation improve rapidly, these protocols can be adapted to rapid changes. Nevertheless, they provide the most important methods for studying basic gene regulatory mechanisms in any type of cell. Here, we discuss some of the pitfalls and facts one has to consider when using these methods.
Cells: Cells must be highly viable and, if using primary isolated cells, purity of cell populations must be guaranteed (e.g., FACS analysis for primary T cells). Even slightly stressed cells can influence the results of these very sensitive sequencing methods and lower the amount of newly transcribed or translated RNA, and lead to unwanted readouts of the stress response in the sequencing results. The centrifugation speed mentioned in this protocol to pellet cells is optimized for primary T cells. Thus, adjust the speed according to the cell type.
4sU effects on cell physiology: In addition to the aforementioned options for verifying minimal perturbation to cellular physiology upon 4sU addition, further and/or additional analysis can be performed, especially when cell numbers are limited. Effects on cell proliferation can be tested by verifying the doubling time of cells by simply counting labeled and unlabeled cells. Nucleolar stress induction could also be tested by analyzing cell morphology via immunofluorescence staining of nucleolin and nuclei. To further verify the impact of 4sU, altered global gene expression could be measured by correlating read counts from labeled total RNA to unlabeled total RNA.
Cell numbers: For in vitro generated T cells, we recommend starting with at least the amount of cells indicated in Table 2. Choose adequate numbers per method according to the type of cells. Since T cells have less cytoplasm and RNA compared to other cells, most likely lower amounts of other cells will be adequate. For ChIP-seq, cell numbers highly depend on the antibody used and the expression level of the protein of interest within the cells. For histone or RNAPII ChIP, fewe cells can be used, whereas cell numbers need to be increased when transcription factors are used, especially if they are expressed at low levels.
4sU-labeling and RNA biotinylation: When using adherent cells, 4sU labeling can be directed as described by Rädle et al.19 Since cells very rapidly incorporate 4sU, it can be added directly to the medium of suspension, adherent, or semi-adherent cells.
It is recommended to start the biotinylation with 60 - 80 µg of RNA. Nevertheless, lower amounts of RNA can be used, although we did not test for less than 30 µg. Add a coprecipitant (e.g., GlycoBlue) when precipitating RNA if the pellet is hard to see. Duffy et al. have also shown that methylthiosulfonate-activated biotin (MTS-biotin) more efficiently reacts with 4sU-labeled RNA than HPDP-biotin31. Hence, it might be worth considering switching to MTS-biotin, in particular, for the recovery of small RNAs, which tend to have fewer uridine residues (refer to the biotinylation protocol mentioned by Duffy et al.; see Purification of 4sU-labeled RNA, Experimental Procedures).
For the recovery of newly transcribed RNA, it is possible to use paramagnetic beads or RNA Cleanup beads of your choice. Always take into account that these kits may or may not purify for specific RNAs. For example, if you are interested in miRNAs, consider using specific kits for miRNA capture and sequencing.
Quantification of newly transcribed RNA: To accurately quantify newly transcribed RNA, measurement should be performed by a suitable fluorometer (see Table of Materials). Within 1 h of 4sU exposure, newly transcribed RNA represents about 1 - 4% of total RNA. Newly transcribed RNA of 1 h labeled, activated T cells consists of ~90 - 94% of rRNA11.
Ribosome Profiling: When establishing the method, we determined that using 1.5x the amount of nuclease than suggested in the original protocol guarantees proper digestion. Also, no adverse effects have been reported for elevated amounts of nuclease. Since it is quite difficult to overdigest the RPFs while they are part of the RNA bound by ribosomal proteins, you can still slightly increase the amount to titrate optimal nuclease digestion.
If less than 500 ng of RPF RNA were recovered in step 4.4.2, repeat rRNA depletion and pool purified RPFs with RNA Clean & Concentrator-5 columns. Alternatively, load two identical samples next to each other on the gel (step 4.5.3) and pool gel slices during RNA elution from the gel (step 4.5.6).
We recommend cutting the RPFs on a gel as tightly as possible to the 28 and 30 nt bands. This helps in eliminating unwanted fragments from rRNAs and tRNAs, which later will become part of your library and reduce sequencing reads for your RPFs.
It is also recommended to avoid UV light during gel purification. This can create nicks in the RNA fragments, as well as pyrimidine dimers, that in the end can severely affect the library preparation and sequencing results.
Library preparation and sequencing of data: Ribosome profiling protocol enables to generate a cDNA library suitable for sequencing. Samples generated by 4sU-labeling can be directly used for library preparation with any appropriate RNA sequencing kit. Since newly transcribed RNA, especially when using short labeling times, may not yet be polyadenylated, no poly-A selection should be performed. Instead, we strongly recommend rRNA depletion to prevent reducing sequencing depth for the actual sample. Using T cells, we started with 400 ng of newly transcribed and total RNA (depending on the kit, see materials), performed rRNA depletion and reduced cycles for PCR amplification to minimize PCR bias. Library preparation can be performed with less starting material. To account for library complexity numbers of PCR cycles should be optimized.
For ChIP-seq there are also many kits available for library preparation. In our hands library preparation worked well starting with 2 ng of ChIP DNA (see materials for a suggestion on which kit to use). Be sure to check the indices for color balance during sequencing. We recommend a sequencing depth of ≥40 x 106 reads each for 4sU-seq, total RNA-seq, and ChIP-seq samples, and ≥80 x 106 reads for ribosome profiling samples. The sequencing depth depends on the sample and the downstream bioinformatics analysis and should be considered carefully. To analyze intronic reads for cotranscriptional splicing, 100 bp paired-end sequencing needs to be chosen.
Sequencing bias: Sequencing has become the gold standard when determining global changes in transcription, translation, or transcription factor binding. Within recent years, existing methods were pushed to their limits or new techniques were developed for sequencing increasingly small starting amounts of RNA. This requires amplification of cDNA, which introduces noise or bias. Recently, unique molecular identifiers (UMIs) were developed to experimentally identify duplicates introduced by PCR. Recently, it was shown that UMIs just mildly improve sequencing power and false discovery rate for differential gene expression32. Nevertheless, consider using unique molecular identifiers (UMIs) for all sequencing libraries to control for library complexity, especially when starting with low amounts of RNA and when many PCR cycles are needed.
Buffer and stock solutions: All buffers for 4sU-seq and ribosome profiling must be prepared under strict RNase-free conditions using nuclease-free water. It is recommended to buy pre-made nuclease-free NaCl, Tris-HCl, EDTA, sodium citrate, and water. To ensure nuclease-free conditions, an RNase decontaminating solution can be used to clean pipettes or surfaces. All buffers for ChIP-seq need to be at least DNase-free and can be stored at room temperature. Always add protease inhibitors and, optionally, phosphatase inhibitors right before use and keep on ice.
Bioinformatics: Analysis of all sequencing data (i.e., ChIP-seq, RNA-seq, and ribsosome profiling) involves quality control (e.g., using FastQC, http://www.bioinformatics.babraham.ac.uk/projects/fastqc/), adapter trimming (e.g., with cutadapt20) followed by mapping to the reference genome for the cells under study. For RNA-seq data (both total and 4sU-seq), as well as ribosome profiling data, a spliced RNA-seq mapper is required, such as ContextMap 221. For ChIP-seq data unspliced alignments, using BWA-MEM22 is sufficient. Gene expression can be calculated using the RPKM model (Reads per Kilobase of Exon per Million Fragments Mapped)1, after determining read counts per gene using a program, e.g. featureCounts23. For peak calling from ChIP-seq data, a number of programs are available, e.g., MACS24 or GEM25. Further downstream analyses can be performed in R26, in particular using tools provided by the Bioconductor project27.
Here, a major challenge in integrating 4sU- and total RNA levels and translational activity from ribosome profiling is normalization. A classical approach to address this problem is to normalize to levels of house-keeping genes. To reduce noise due to random fluctuations for individual house-keeping genes, it is recommended to not just use a few house-keeping genes but median levels for a larger set, e.g. the >3,000 house-keeping genes compiled by Eisenberg and Levanon28. For calculation of RNA turnover rates from ratios of 4sU- to total RNA, normalization is based on median turnover rates (e.g., assuming an RNA half-life of 5h)29. However, since this assumes no overall changes for house-keeping genes, we recommend using analysis approaches independent of normalization, e.g., correlation-based clustering of a time-series of the different data types to identify groups of genes with distinct behavior in transcription and translation during activation. For a detailed description on bioinformatical integration of the different data types, we refer to the original publication11.
Analysis of turnover rates and data integration: A recently published paper33 comparing half-lives determined by a multiplexed gene control (MGC) to global methods could show that half-lives correlated best with those obtained by metabolic labeling methods compared to other methods (e.g., general inhibition of transcription by drugs). However, it should be mentioned that differences between half-life calculations can arise and have been described 15,34. We account for most of the problems and differences that are introduced by the stress response due to prolonged 4sU exposure. Therefore, it is indispensable to exclude the stress response introduced by 4sU-labeling. To further validate turnover rates, we recommend the use of MGCs.
Additionally, a data set as generated here could also be used for a more integrative data analysis (e.g., regulation of long non-coding RNAs)35,36.