$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
The described Capture Hi-C protocol is based on the preparation of the genome-wide 3C-template using a four-base cutter (DpnII). The subsequent enrichment of ligation fragments across the genomic region of interest is obtained by hybridization of an array of tiling RNA probes and their streptavidin-based capture according to the target enrichment system used in this study (Figure 1). Biotinylated RNA probes were selected as they show tighter binding affinity to their targets compared to DNA probes52,60. Captured libraries are then indexed and pooled for multiplexed high-throughput sequencing. Capture Hi-C data can be visualized as high-resolution Hi-C interaction maps but also as 4C-like single view-point contact maps to specifically visualize the interactions of smaller sequences such as promoters or enhancers within the entire captured region. The workflow of the protocol is shown in Figure 4. Pre-sequencing quality controls are shown in Figure 2 and include the assessment of proper digestion and re-ligation of the 3C template and its efficient shearing and purification across the different steps of the protocol. The sheared 3C template DNA is expected to run between 150 to 700 bp, and no enrichment of fragments >2 kb should be detected. During the following steps, several bead-based DNA cleanup and size selection steps are performed, first after the shearing, then after the pre-capture and post-capture PCRs. Cleaned libraries show a distinct fragment enrichment profile as visualized on a high sensitivity DNA bioanalyzer (Figure 2). The mean fragment size increases over the course of the library preparation due to the ligation of adaptors, sequencing, and indexing primers. Post-sequencing quality controls are obtained via Hi-C Pro and shown in Figure 3. Many different bioinformatics software applications have been proposed for 3C-like data processing and analysis. Among them, the HiC-Pro pipeline is one of the most popular solutions, allowing the processing of raw sequencing data to the final contact maps at various resolutions55. HiC-Pro uses a two-step mapping strategy to align the sequencing reads on the reference genome. The 3C products are then reconstructed and filtered out to remove non-informative pairs of contact and to generate the contact maps. In addition, it is able to use a list of known polymorphisms to perform allele-specific analysis and to separate the contacts coming from the two parental alleles in distinct contact maps. More recently, HiC-Pro has been included and extended into the nf-core framework (nf-core-hic), providing a highly scalable and reproducible community-driven pipeline61,62.
To capture the mouse Xic, an array of 28,913 RNA probes tiling 3 Mb of the X chromosome was designed. This region includes the key player in XCI, the long noncoding gene Xist, and its known ~800 kb regulatory landscape (Figure 5). This ~800 kb region is partitioned into two TADs: one including the Xist promoter and its known positive regulators (i.e., the noncoding transcripts Ftx, Jpx, and Xert and the protein-coding gene Rnf12), and the neighboring TAD encompassing the negative cis-regulators of Xist (i.e., its antisense transcript Tsix, the enhancer element Xite, and the noncoding transcript Linx) (for review44,45).
By applying the described Capture Hi-C protocol to the Xic, the topological organization of this locus was obtained at unprecedented resolution (Figure 6 and Figure 7). This is particularly clear when comparing the Capture Hi-C profile to previously published 5C47 (Figure 6 and Figure 7; Supplementary Table 1) and Hi-C61 (Figure 6 and Figure 7; Supplementary Table 1) profiles. For instance, sub-TAD structures are more evident — the TAD containing the Xist promoter (Xist-TAD) is clearly subdivided into two smaller domains (Figure 6A, blue arrowhead). Previously, this could only be visually "guessed" from the 5C profile (Figure 6B), albeit the detection of a boundary in this region using the insulation score algorithm. Likewise, the resolution of the Capture Hi-C profile allows the identification of two smaller domains in the neighboring TAD (Figure 6A, B), which contains the promoter of the Tsix locus (Tsix-TAD); this was not previously achieved with 5C (Figure 6B). Of note, topological boundaries determined by the insulation score from the Capture Hi-C and 5C data are generally detected at slightly different locations and with different relative strengths.
Moreover, other sub-TAD structures such as contact loops are clearly visible from the Capture Hi-C data, such as the loop between Xist and Ftx (Figure 7A), previously identified with Capture-C63, and the loop between Xist and Xert (Figure 7B), recently identified using a similar protocol for Capture Hi-C48. Other contacts can also be mapped more precisely due to the increased resolution of the Capture Hi-C profiles, such as those forming the known contact hotspots within the Tsix-TAD between the Linx, Chic1, and Xite loci (Figure 7A).
In comparison with the Hi-C data shown in Figure 7, Capture Hi-C allowed for a fourfold increase in resolution, yet it required only one-fourth of sequencing depth (i.e., 126 M reads versus 571 M) (Supplementary Table 1). This increase in resolution allows for the detection of subTADs and looping interactions that could not be detected by Hi-C at the sequencing depth shown in Figure 6 and Figure 7. The described protocol for Capture Hi-C thus allows for a much more detailed, high-resolution characterization of a large genomic region of interest, when compared to previous approaches.

Figure 1: Probe design. Schematic representation of the strategy used for probe design. Regions of 300 bp upstream and downstream of each DpnII restriction site across the 3 Mb target region were selected and tiled with overlapping biotinylated RNA probes. One of these selected regions is shown, chrX: 102,474,805-102,475,500. No more than 40 bases of repetitive sequences are allowed in each probe. Please click here to view a larger version of this figure.

Figure 2: Capture Hi-C pre-sequencing quality controls. (A) Representative example of 3C template quality controls. 200 ng of DNA were loaded on a 1% agarose gel. Lane 1: 1 kb ladder. Lane 2: Undigested, cross-linked, and intact chromatin runs as a sharp band at >10 kb. Lane 3: DpnII-digested cross-linked chromatin runs as a smear between 1 kb to 3 kb in size. Lane 4: Final 3C library or template; free ends of digested cross-linked DNA fragments are re-ligated. The DNA smear of lower molecular size is almost undetectable, and the ligation product is detected as a band of >10 kb. (B) Representative examples of high sensitivity bioanalyzer DNA profiles. Top left: successfully sheared 3C library showing a distribution of fragment size between 150 bp and 700 bp. Top right: unsatisfactory sheared 3C library. Unsheared DNA is detected as broad enrichment of fragments >2 kb. (C) Bottom left: sheared DNA sample following a 1:1 left side-size selection using SPRI beads. Fragments of ~300 bp are enriched. Bottom middle: Pre-capture PCR profile after ligation of paired-end adaptors according to the manufacturer's protocol. Bottom right: final Capture Hi-C library including adaptors, sequencing, and indexing primers for multiplexed sequencing. Abbreviations: bp = base pairs, FU = arbitrary fluorescence unit. Please click here to view a larger version of this figure.

Figure 3: Capture Hi-C post-sequencing quality controls with HiC-Pro. (A) Example of mapping rate on the reference genome for the first mate of the sequencing pairs. The light blue fraction represents the reads aligned by HiC-Pro and spanning a ligation junction. This metric can thus be used to validate the experimental ligation step. (B) Once sequencing mates are aligned on the genome, only uniquely aligned read pairs are kept for analysis. (C) Non-valid pairs (in red) such as dangling-end, self-circle, or re-ligation are discarded from the analysis. The fraction of valid pairs is a good indicator of the ligation and pull-down efficiency. (D) The valid pairs can be further divided into intra/inter-chromosomal and short/long-range contacts. Duplicated read pairs that are likely to represent PCR artifacts are discarded from the analysis. (E) For allele-specific analysis, HiC-Pro reports the number of allelic reads supported by either one or two mates for each parental genome (i.e., C57BL/6J x CASTEi/J). The same fraction of reads assigned to the maternal and paternal allele are expected. (F) Finally, only valid pairs overlapping the capture region are selected to build the contact maps. Capture-capture pairs represent contacts within the targeted region, while capture-reporter pairs involve interaction between the targeted region and an off-target one. Please click here to view a larger version of this figure.

Figure 4: Workflow of Capture Hi-C protocol. Schematic representation of different protocol steps. To generate the genome-wide 3C template, chromatin is first cross-linked with formaldehyde and then digested with the DpnII restriction enzyme. Free DNA ends are then re-ligated, cross-linking is reversed, and DNA is purified. To enrich fragments encompassing the target region, an array of biotinylated RNA probes is hybridized to the 3C template and captured by streptavidin-mediated pull-down. Capture libraries are processed for multiplexed sequencing, and valid ligation fragments are quantified to infer the frequency of chromatin contacts across the target, which are visualized as high-resolution interaction maps. Please click here to view a larger version of this figure.

Figure 5: Overview of the region encompassing the Xic on the mouse X chromosome. Schematic representation of the mouse X chromosome and zoom in of the 3 Mb captured region (ChrX: 102,475,000-105,475,000). The targeted region includes ~800 kb of DNA corresponding to the Xic, the master regulatory locus of XCI. The Xic includes the long noncoding genes, Xist, a key player of XCI, and its regulatory landscape. Positive regulators of Xist are shown in green, and negative regulators in purple. Please click here to view a larger version of this figure.

Figure 6: Capture Hi-C, 5C, and Hi-C interaction maps across the 3 Mb captured region. (A) Capture Hi-C interaction map of the 3 Mb target encompassing the mouse Xic at 10 kb resolution (this study). (B) 5C interaction map of the same target region as in A at 6 kb resolution (data reprocessed from47). Repetitive regions not included in the analyses are masked in white. The 5C data require their own bioinformatics processing (see47). After cleaning and alignment, the 5C maps at the primer resolution are binned using a running median (window = 30 kb, step = 5) to reach a final resolution of 6 kb. (C) Hi-C interaction map of the same genomic region as in A and B at 40 kb resolution (data reprocessed from64). All interaction maps were generated from mouse ESCs. The insulation score was calculated using cooltools and is represented as histograms with insulation minimas at TAD boundaries. TAD boundaries are shown as vertical lines below the map. The height of each line indicates boundary strength. Genes are shown as arrows pointing in the direction of transcription. Sub-TAD boundaries that are detected exclusively or more precisely in Capture Hi-C maps are indicated by magenta and blue arrowheads for sub-TADs in the Tsix and Xist TADs, respectively. Please click here to view a larger version of this figure.

Figure 7: Capture Hi-C, 5C, and Hi-C interaction maps across 1 Mb within the captured region. (A) Capture Hi-C interaction map of the 1 Mb genomic region encompassing the mouse Xic at 5 kb resolution (this study). (B) 5C interaction map of the same genomic region as in A. at 6 kb resolution (data reprocessed from47). Repetitive regions not included in the analyses are masked in white. Of note, the 5C data require their own bioinformatics processing (see47). After cleaning and alignment, the 5C maps at the primer resolution are binned using a running median (window = 30 kb, step = 5) to reach a final resolution of 6 kb. (C) Hi-C interaction map of the same genomic region as in A and B of Hi-C at 20 kb resolution (data reprocessed from64). All interaction maps were generated from mESCs. The insulation score was calculated using cooltools and is represented as histograms with insulation minimas at TAD boundaries. TAD boundaries are shown as vertical lines below the map. The height of each line indicates boundary strength. Genes are shown as arrows pointing to the direction of transcription. Contact loops that are detected exclusively or more precisely in Capture Hi-C are indicated by magenta and blue asterisks for loops in the Tsix and Xist TADs, respectively. Please click here to view a larger version of this figure.
Supplementary Table 1: Post-sequencing statistics for the datasets used in this manuscript: Capture Hi-C (this study), Hi-C64, and 5C47. Please click here to download this File.