In human complex disease, genetic variants contribute to disease susceptibility and quantifying these variants may be useful for understanding pathogenesis, identifying high risk patient groups, and treatment responders. Indeed, the promise of precision medicine is dependent upon utilizing genomic information to identify patient subgroups1. Unfortunately, within the complex disease biology space, where disease phenotypes are underpinned by substantial genetic heterogeneity, low penetrance, and variable expressivity, cohort size requirements for genome-wide approaches to identify novel candidates are often prohibitively large2. Alternatively, a targeted candidate gene approach begins with an a priori hypothesis about specific genes/pathways in disease etiology3. Pathway analysis tools are commonly used to investigate the pathophysiology of an identified target loci, generating numerous candidate pathways to be explored. We demonstrate here a multiplexed genotyping approach allowing for the investigation of tens to hundreds of SNPs with one assay, suited to human cohort studies4. This approach is relatively high through-put, permitting hundreds to thousands of DNA samples to be genotyped for novel discovery studies and investigation of specific pathways. The methods outlined here are useful for identifying risk alleles and their associations with clinical traits in a relatively rapid and inexpensive manner. This platform has been highly advantageous for screening and diagnostic purposes5,6, and more recently, for microbial infection7 and human papillomavirus8.
This protocol begins with selection of a set of genes for investigation, i.e., the target regions, typically determined through literature searching, or a priori hypotheses for involvement in the disease process; or perhaps selected for replication as the leading associations of a discovery genome-wide association (GWA) study. From the gene set, the researcher will select a refined list of tag SNPs. That is, the linkage disequilibrium (LD), or correlation, amongst variants in the region is used to identify a representative 'tag SNP' for a group of SNPs in high LD, known as a haplotype. The high LD of the region means that the SNPs are often inherited together such that genotyping one SNP is sufficient to represent the variation at all SNPs in the haplotype. Alternatively, if following up on a definitive list of SNPs from many regions, e.g., replication for a GWA study, this process may be unnecessary. For multiplexed genotyping, an assay must then be designed around these targets such that the amplification primers are of differing mass to those of the extension primers and products to produce interpretable mass spectra. These parameters are easily implemented by a multiplexed genotyping assay design tool. The forward and reverse primers from this design will be used to target the markers of interest and amplify the sequence containing the SNP. The extension primers attach directly proximal to the SNPs and a single, mass-modified, 'terminator' base that is complementary to the SNP is added. The terminator base prevents further extension of the DNA. The mass-modification of the base enables fragments differing by a single base to be detected by mass spectrometry. The plate containing the genotyping chemistry is then applied to a chip for measurement on a mass spectrometry platform. After applying appropriate quality controls to the raw genotyping calls detected by the system, the data can be exported and used for statistical analysis to test for association with disease phenotypes.