Modular design of Promoter Capture Hi-C
Promoter Capture Hi-C is designed to specifically enrich Hi-C libraries for interactions involving promoters. These interactions comprise only a subset of ligation products present in a Hi-C library.
Capture Hi-C can easily be modified to enrich Hi-C libraries for any genomic region or regions of interest by changing the capture system. Capture regions can be continuous genomic segments44,45,46,48, enhancers that have been identified in PCHi-C ('Reverse Capture Hi-C'35), or DNase I hypersensitive sites49. The size of the capture system can be adjusted depending on the experimental scope. For example, Dryden et al. target 519 bait fragments in three gene deserts associated with breast cancer44. The capture system by Martin et al. targets both continuous genomic segments ('Region Capture': 211 genomic regions in total; 2,131 restriction fragments) and selected promoters (3,857 gene promoters)45.
SureSelect libraries are available in different size ranges: 1 kb to 499 kb (5,190–4,806), 500 kb to 2.9 Mb (5,190–4,816), and 3 Mb to 5.9 Mb (5,190–4,831). As each individual capture biotin-RNA is 120 nucleotides long, these capture systems accommodate a maximum of 4,158, 24,166 and 49,166 individual capture probes, respectively. This corresponds to 2,079, 12,083, and 24,583 targeted restriction fragments, respectively (note that the numbers for restriction fragments are lower bounds based on the assumption that two individual capture probes can be designed for every restriction fragment — in reality due to repetitive sequences this will not be the case for every restriction fragment (see also Figure 1B, C), resulting in a higher number of targetable restriction fragments for a constant number of available capture probes).
The protocol described here is based on the use of a restriction enzyme with a 6 bp recognition site to uncover long-range interactions. Using a restriction enzyme with a 4 bp recognition site for greater resolution of more proximal interactions is also possible40,49.
Limitations of PCHi-C
One inherent limitation of all chromosome conformation capture assays is that their resolution is determined by the restriction enzyme used for the library generation. Interactions that occur between DNA elements located on the same restriction fragment are invisible to 'C-type' assays. Further, in PCHi-C, in some cases more than one transcription start site can be located on the same promoter-containing restriction fragment, and PIRs in some cases harbor both active and repressive histone marks, making it difficult to pinpoint which regulatory elements mediate the interactions, and to predict the regulatory output of promoter interactions. Using restriction enzymes with 4 bp recognition sites mitigates this issue but comes at the expense of vastly increased Hi-C library complexity (Hi-C libraries generated with 4 bp recognition site restriction enzymes are at least 100 times more complex than Hi-C libraries generated with 6 bp recognition site restriction enzymes), and the associated costs for next generation sequencing.
Another limitation is that the current PCHi-C protocol requires millions of cells as starting material, precluding the analysis of promoter interactions in rare cell types. A modified version of PCHi-C to enable the interrogation of promoter contacts in cell populations with 10,000 to 100,000 cells (for example cells during early embryonic development or hematopoietic stem cells) would therefore be a valuable addition to the Capture Hi-C toolbox.
Finally, like all methods that rely on formaldehyde fixation, PCHi-C only records interactions that are 'frozen' at the time point of fixation. Thus, to study the kinetics and dynamics of promoter interactions, methods such as super-resolution live cell microscopy are required alongside PCHi-C.
Methods to dissect spatial chromosome organization at high resolution
The vast complexity of chromosomal interaction libraries prohibits the reliable identification of interaction products between two specific restriction fragments with statistical significance. To circumvent this problem, sequence capture has been used to enrich either Hi-C33,34,40,44 or 3C50,51 libraries for specific interactions. The major advantage of using Hi-C libraries over 3C libraries for the enrichment step is that Hi-C, unlike 3C, includes an enrichment step for genuine ligation products. As a consequence, the percentage of valid reads in PCHi-C libraries is approximately 10-fold higher than in Capture-C libraries50, which contained around 5–8% valid reads after HiCUP filtering. Sahlen et al. have directly compared Capture-C to HiCap, which like PCHi-C uses Hi-C libraries for capture enrichment, in contrast to Capture-C which uses 3C libraries. Consistent with our findings, they found that Capture-C libraries are mainly composed of un-ligated fragments40. In addition, HiCap libraries had a higher complexity than Capture-C libraries40.
A variant of Capture-C, called next-generation Capture-C52 (NG Capture-C) uses one oligo per restriction fragment end, as previously established in PCHi-C33,34, instead of overlapping probes used in the original Capture-C protocol50. This increases the percentage of valid reads compared to Capture-C modestly, but NG Capture-C employs two sequential rounds of capture enrichment, and a relatively high number of PCR cycles (20 to 24 cycles in total, compared to 11 cycles typically for PCHi-C), which inevitably results in higher numbers of sequence duplicates and lower library complexity. In trial experiments during the optimization of PCHi-C, we found that the percentage of unique (i.e., not duplicated) read pairs was only around 15% when we used 19 PCR cycles (13 cycles pre-capture + 6 cycles post-capture; data not shown), however optimization to a lower number of PCR cycles, typically yields 75–90% unique read pairs. Thus, reducing the number of PCR cycles substantially increases the amount of informative sequence data.
A recent method combines ChIP with Hi-C to focus on chromosomal interactions mediated by a specific protein of interest (HiChIP53). Compared to ChIA-PET54, which is based on a similar rationale, HiChIP data contains a higher number of informative sequence reads, allowing for higher-confidence interaction calling53. It will be very interesting to directly compare the corresponding HiChIP and Capture Hi-C data sets once they become available (for example HiChIP using an antibody against the cohesin unit Smc1a53 with Capture Hi-C for all Smc1a bound restriction fragments) side by side. One inherent difference between these two approaches is that Capture Hi-C does not rely on chromatin immunoprecipitation, and therefore is capable of interrogating chromosomal interactions irrespective of protein occupancy. This enables comparison of 3D genome organization in the presence or absence of specific factor binding, as has been used to identify PRC1 as a key regulator of mouse ESC spatial genome architecture7.
PCHi-C and GWAS
Genome-wide association studies (GWAS) have revealed that greater than 95% of disease-associated sequence variants are located in non-coding regions of the genome, often at great distances to protein-coding genes55. GWAS variants are often found in close proximity to DNase I hypersensitive sites, which is a hallmark of sequences with potential regulatory activity. PCHi-C and Capture Hi-C have been used extensively to link promoters to GWAS risk loci implicated in breast cancer44, colorectal cancer48, and autoimmune disease35,45,46. A PCHi-C study on 17 different human hematopoietic cell types found SNPs associated with autoimmune disease were enriched in PIRs in lymphoid cells, whereas sequence variants associated with platelet and red blood cell specific traits were predominantly found in the macrophages and erythroblasts, respectively35,56. Thus, tissue-type specific promoter interactomes uncovered by PCHi-C may help to understand the function of non-coding disease-associated sequence variants and identify new potential disease genes for therapeutic intervention.
Characteristics of promoter-interacting regions
Several lines of evidence link promoter interactomes to gene expression control. First, several PCHi-C studies have demonstrated that genomic regions interacting with promoters of (highly) expressed genes are enriched in marks associated with enhancer activity, such as H3K27 acetylation and p300 binding33,34,37. We found a positive correlation between gene expression level and the number of interacting enhancers, suggesting that additive effects of enhancers result in increased gene expression levels34,35. Second, naturally occurring expression quantitative trait loci (eQTLs) are enriched in PIRs that are connected to the same genes whose expression is affected by the eQTLs35. Third, by integrating TRIP57 and PCHi-C data, Cairns et al. found that TRIP reporter genes mapping to PIRs in mouse ESCs show stronger reporter gene expression than reporter genes at integration sites in non-promoter-interacting regions58, indicating that PIRs possess transcriptional regulatory activity. Together, these findings suggest that promoter interactomes uncovered by PCHi-C in various mouse and human cell types include key regulatory modules for gene expression control.
It is worth noting that enhancers represent only a small fraction (~20%) of all PIRs uncovered by PCHi-C33,34. Other PIRs could have structural or topological roles rather than direct transcriptional regulatory functions. However, there is also evidence that PCHi-C may uncover DNA elements with regulatory function that do not harbor classical enhancer marks. In a human lymphoid cell line, the BRD7 promoter was found to interact with a region devoid of enhancer marks that was shown to possess enhancer activity in reporter gene assays33. Regulatory elements with similar characteristics may be more abundant than currently appreciated. For example, a CRISPR-based screen for regulatory DNA elements identified unmarked regulatory elements (UREs) that control gene expression but are devoid of enhancer marks59.
In other cases, PIRs have been shown to harbor chromatin marks associated with transcriptional repression. PIRs and interacting promoters bound by PRC1 in mouse ESCs were engaged in an extensive spatial network of repressed genes bearing the repressive mark H3K27me37. In human lymphoblastoid cells, a distant element interacting with the BCL6 promoter repressed transgene reporter gene expression33, suggesting that it may function to repress BCL6 transcription in its native context.
PIRs enriched for occupancy of the chromatin insulator protein CTCF in human ESCs and NECs37 may represent yet another class of PIRs. Collectively, these results suggest that PIRs harbor a collection of gene regulatory activities yet to be functionally characterized.