For many years there has been an unmet need to identify the set of genes bound and regulated by a given protein genome wide, in particular those in the transcription factor class.
Odom et al.1 used chromatin immunoprecipitation (ChIP) combined with promoter microarrays to systematically identify the genes occupied by pre-specified transcriptional regulators in human liver and pancreatic islets. Subsequently, Johnson et al.2 developed a large-scale chromatin immunoprecipitation assay based on direct ultra high-throughput DNA sequencing (ChIP-seq) in order to comprehensively map protein-DNA interactions across entire mammalian genomes. As a test case, they mapped in vivo the binding of the neuron-restrictive silencer factor (NRSF) to 1946 locations in the human genome. The data displayed sharp resolution of binding position (+50 base pairs), which facilitated both the isolation of motifs and the identification of NRSF-binding motifs. These ChIP-seq data also had high sensitivity and specificity and statistical confidence (P < 10−4), properties that are important for inferring new candidate interactions.
Robertson et al.3 also used ChIP-seq in order to map STAT1 targets in interferon-γ (IFN-γ)-stimulated and unstimulated human HeLa S3 cells in vivo. By ChIP-seq, using 15.1 and 12.9 million uniquely mapped sequence reads, and an estimated false discovery rate of less than 0.001, they identified 41,582 and 11,004 putative STAT1-binding regions in stimulated and unstimulated cells, respectively. Of the 34 loci known to contain STAT1 interferon-responsive binding sites4-8, ChIP-seq found 24 (71%). ChIP-seq targets were enriched in sequences similar to known STAT1 binding motifs. Comparisons with two existing ChIP-PCR data sets suggested that ChIP-seq sensitivity was between 70% and 92% and specificity was at least 95%. Additionally, it was clear that ChIP-seq offers both low analytical complexity and sensitivity that increases with sequencing depth.
As such, "next-generation" genome sequencing technologies provide 1-2 orders of magnitude increase in the amount of sequence that can be cost-effectively generated over older technologies9. ChIP-seq methods therefore directly provide whole-genome coverage for effective profiling of mammalian protein-DNA interactions3.
In 2006, a strong association of variants in the transcription factor 7-like 2 (TCF7L2) gene with type 2 diabetes was discovered10. Other investigators have already independently replicated this finding in different ethnicities and, interestingly, from the first genome wide association studies of type 2 diabetes published in Nature11,12, Science13-15 and elsewhere16,17, the strongest association was indeed with TCF7L2; this is now considered the most significant genetic finding in type 2 diabetes to date18-20. In addition, TCF7L2 has been linked to cancer risk 21,22; indeed, this connection became more obvious when the 8q24 locus revealed by genome wide association studies of a number of cancers, including colorectal carcinomas, was shown to be due to an extreme upstream TCF7L2-binding element driving the transcription of MYC23,24. As such, there is great interest in determining the downstream genes regulated by this key transcription factor.
Based on experience with TCF7L2 as an example of the methodology, this paper outlines how to generate high quality ChIP DNA template. ChIP was carried out in the colorectal carcinoma cell line, HCT116, for subsequent sequencing in order build a high-resolution map of the genes bound by TCF7L225 in an endeavor to yield further insight in to its key role in the pathogenesis of complex traits.