$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
During the past decade, there has seen a dramatic improvement in sequencing technologies, allowing the study of a wide range of genomes in large numbers of samples, and with astonishing resolution. The Encyclopedia of DNA Elements (ENCODE) Consortium, a large-scale multi-institutional effort spearheaded by the National Human Genome Research Institute of the National Institutes of Health, has provided insights into how individual transcription factors and other regulatory proteins bind to and interact with the genome. The initial effort characterized specific DNA-protein interactions, as assessed by Chromatin immunoprecipitation (ChIP) for over 100 known DNA-binding proteins1. Alternative methods such as DNase footprinting2 and formaldehyde assisted isolation of regulatory elements (FAIRE)3 have also been used to locate specific regions of the genome interacting with proteins, but with the obvious limitation that these experimental approaches do not identify the interacting proteins. Despite the extensive efforts over the past years, no technology has emerged that efficiently allows the comprehensive characterization of protein-DNA interactions in chromatin, and the identification and quantification of chromatin-associated proteins.
To address this need, we developed a novel approach which we termed as Hybridization Capture of Chromatin-Associated Proteins for Proteomics (HyCCAPP). Initially developed in yeast4,5,6, the approach isolates crosslinked chromatin regions of interest (with bound proteins) using sequence-specific hybridization capture. After isolation of the protein-DNA complexes, approaches such as mass spectrometry can be used to characterize the set of proteins bound to the sequence of interest. Thus, HyCCAPP can be considered as a non-biased approach to uncover novel DNA-protein interactions, in the sense that it does not rely on antibodies and it is completely agnostic about the proteins that might be found. There are other approaches capable to uncover novel DNA-interacting proteins7, but most rely on ChIP-like methods8,9,10, plasmid insertions11,12,13,14, or regions with high copy numbers15. In contrast, HyCCAPP can be applied to multi- and single-copy regions, and it does not require any prior information about the proteins in the region. In addition, while some of the methods mentioned above have valuable features, notably avoiding the need for DNA-protein crosslinking reactions, the unique feature of HyCCAPP is that it can be applied to single-copy regions in unmodified cells, and without any prior knowledge about putative binding proteins, or available antibodies.
At this point, HyCCAPP has predominantly been applied to the analysis of various genomic regions in yeast4,5,6, and was recently used to analyze protein-DNA interactions in alpha-satellite DNA, a repeat region in the human genome16. As part of our ongoing work, we have adapted the hybridization capture approach initially developed for yeast chromatin to be applicable to the analysis of human cells, and present here a modified protocol that allows the selective capture of single-copy target regions in the human genome with efficiencies similar to our initial studies in yeast. This new optimized protocol now allows the adaptation and utilization of the technology to interrogate protein-DNA interactions across the human genome, using mass spectrometry or other analytical approaches.
It is important to emphasize that the HyCCAPP method is meant for the analysis of specific target regions and is not yet suitable for genome-wide analyses. The technology is especially useful when dealing with regions for which there is scarce information about interacting proteins, or when a more comprehensive in-depth analysis of interacting proteins at a specific genome locus is desired. HyCCAPP is meant to uncover DNA-binding proteins but not characterize accurately the specific protein binding sites in genomic DNA. In its current implementation, the methodology does not provide information about the DNA binding sequences or motifs for individual proteins. Therefore, it nicely complements existing technologies such as FAIRE, and may allow the identification of novel binding proteins in genomic regions identified by an initial FAIRE analysis.