Oncology care has been transformed in recent years by the use of next-generation sequencing (NGS) technologies to inform diagnosis, prognosis, and therapy selection1. NGS testing has become the standard of care in multiple tumor types for 1) establishing a diagnosis of a molecularly defined tumor, 2) using tumor-type agnostic biomarker results for therapy decisions, and 3) determining risk stratification for specific tumor types2,3. Accurate and timely results from NGS and other biomarker testing have become a critical component for patient care and planning in oncology.
NGS testing is typically performed in either large reference laboratories or in hospital-based/academic laboratories, when institutions have sufficient expertise and resources to develop and maintain such testing. The Genomics Organization for Academic Laboratories (GOAL) organization was conceived as a collaborative network, focused on the sharing of academic knowledge and reagent costs4. A significant effort of the GOAL consortium has involved training sessions for critical components for the review and reporting of clinical studies. These have been developed as practical instructional sessions and include an interactive workshop on the integrative genomics viewer (IGV) to demonstrate its basic and advanced features in genomic medicine.
IGV is a data visualization tool designed to allow users to easily review NGS data and other large datasets, given that most NGS data are vast and unwieldy without a graphical user interface5. IGV was not initially designed for clinical variant analysis. However, given its ease of use for all molecular professionals, regardless of their level of bioinformatics experience, it has been widely adopted by many laboratories as part of the complex data review workflow6. This study was assembled as a mechanism to disseminate this information to the broader community of genomic laboratory professionals.
The clinical need for using IGV as part of a clinical workflow includes detecting sequencing errors, handling complexities in genomic variant calling, and other recurrent alterations where additional visualization can assist in appropriate data interpretation7. The use of IGV requires standard bioinformatic files originating from an NGS study8. Core features of the IGV software allow visualization of sequencing data and require configuring a set of preferences and options for clinical variant interpretation (Supplementary File 1).
The short-read sequencing workflow is typified by FASTQ, binary alignment map (BAM), and variant call format (VCF) file standards. NGS achieves its high throughput by generating millions of short reads in parallel. While this strategy greatly accelerates data production, each read is only a fragment of the original DNA or RNA. It is affixed to some position on an NGS flow cell by its adapter sequence. Sequence and base quality information from optical processing is recorded in FASTQ files. These raw reads are nucleotide sequences along with their per-base quality scores. Therefore, an early step in processing raw reads is to align each read to a reference genome8. As each read is assigned a specific location in the reference genome with the highest probability of alignment, the reads that correspond to the same location "pile up" at each locus, and such data is stored in a BAM file following the sequence alignment mapping (SAM) specifications (available at https://samtools.github.io/hts-specs/SAMv1.pdf; last verified: 9/30/2025) (Figure 1A). This information can be visually inspected by a genome browser such as IGV. As illustrated in Figure 1B, comparing the reads to one another or to the reference sequence itself can indicate differences or discrepancies between the nucleotide sequences. Summaries of only these "variant" positions across the entire genome are compiled into a VCF file, which can therefore be much smaller in size than a BAM file but still contain much of the useful information required for variant review. For that reason, the VCF file is commonly the final output of variant-calling pipelines (Figure 1C). Thus, while FASTQ files contain the raw read and quality data, the BAM file provides insight into how confidently those reads map to the genome, and the VCF represents the differences between the reference sequence and the sample sequence, providing a much more manageable amount of data for review for potential reporting.
The six clinical vignettes in this study illustrate common analytical scenarios associated with somatic variant detection in molecular laboratories as well as the problems that complicate the correct understanding and interpretation of the underlying data. Overall, these vignettes were chosen to illustrate increasingly sophisticated interpretative scenarios and how utilizing advanced features in IGV can better describe comprehensive genomic features. Although this study focuses on somatic variant interpretation, these features are applicable for utilizing IGV in the visualization and interpretation of germline alterations6. Of note, the terms used in this article are based on sequencing-by-synthesis chemistry and, where appropriate, other analogous terms/concepts for other platforms.
Each clinical vignette begins with a very brief clinical context, followed by instructions on how to investigate the variant or variants and resolve the question: which (if any) variant present is real and which (if any) is an artifact. The clinical context, along with the particulars of the sequencing data, is important for answering that question.