$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
The identification of regulatory elements for gene expression has been facilitated by the Encyclopedia of DNA Elements (ENCODE) Project that comprehensively annotated functional activity for 80% of the human genome1,2. The identification of the sites for in vivo transcription factor binding, DNaseI hypersensitivity, and epigenetic histone and DNA methylation modifications in individual cell types paved the way for the functional analyses of candidate regulatory elements for target gene expression. Armed with these findings, we are faced with the challenge of determining the functional interconnectivity between regulatory elements and genes. Specifically, what is the relationship between a given target gene and its enhancer(s)? The chromatin conformation capture (3C) method directly addresses this question by identifying physical, and likely functional, interactions between a region of interest and candidate interacting sequences through captured events in fixed chromatin3. As our understanding of chromatin interactions has increased, however, it is clear that the investigation of preselected candidate loci is insufficient to provide a complete understanding of gene-enhancer interactions. For example, ENCODE used the high-throughput chromosomal conformation capture carbon copy (5C) method to examine a small portion of the human genome (1%, pilot set of 44 loci) and reported complex interconnectivity of the loci. Genes and enhancers with identified interactions averaged 2–4 different interacting partners, many of which were hundreds of kilobases away in linear space4. Further, Li et al. used Chromatin Interaction Analysis by Paired-End Tag Sequencing (ChIA-PET) to analyze whole-genome promoter interactions and found that 65% of RNA polymerase II binding sites were involved in chromatin interactions. Some of these interactions resulted in large, multi-gene complexes spanning hundreds of kilobases of genomic distance and containing, on average, 8-9 genes each5. Together, these findings highlight the need for unbiased whole-genome methods for interrogating chromatin interactions. Some of these methods are reviewed in Schmitt et al.6.
More recent methods for chromatin conformation capture studies coupled with next-generation sequencing (Hi-C and 4C-seq) enable the discovery of unknown sequences interacting with a region of interest6. Specifically, circular chromosome conformation capture with next-generation sequencing (4C-seq) was developed to identify loci interacting with a sequence of interest in an unbiased manner7 by sequencing DNA from captured chromatin proximal to the region of interest in 3D space. Briefly, chromatin is fixed to preserve its native protein-DNA interactions, cleaved with a restriction enzyme, and subsequently ligated under dilute conditions to capture biologically relevant "tangles" of interacting loci (Figure 1). The cross links are reversed to remove the protein, thus leaving the DNA available for additional cleavage with a second restriction enzyme. A final ligation generates smaller circles of interacting loci. Primers to the sequence of interest are then used to generate an amplified library of unknown sequences from the circularized fragments, followed by downstream next generation sequencing.
The protocol presented here, which focuses on sample preparation, makes two major alterations to existing 4C-seq methods8,9,10,11,12. First, it uses a qPCR-based method to empirically determine the optimal number of amplification cycles for 4C-seq library preparation steps and thus mitigates the potential for PCR bias stemming from over-amplification of libraries. Second, it uses an additional restriction digest step in an effort to reduce the uniformity of known "bait" sequences that hinders accurate base-calling by the sequencing instrument and, hence, maximizes the unique, informative sequence in each read. Other protocols circumvent this issue by pooling many (12-15)8 4C-seq libraries with different bait sequences and/or restriction sites, a volume of experiments which may not be achievable by other laboratories. The modifications presented here allow a small number of experiments, samples, and/or replicates to be indexed and pooled into a single lane.