Method Article

Comprehensive Evaluation of Genotype Imputation Tools for Ultra-low-depth Whole-Genome Sequencing Data

DOI:

10.3791/68879

December 12th, 2025

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Three imputation tools-STITCH, QUILT2, and GLIMPSE2-were benchmarked across varying sequencing depths and sample sizes, using CKB and EAS reference panels. The results provide a practical framework for selecting appropriate imputation strategies in ultra-low-depth sequencing data, facilitating large-scale population genomic and complex trait studies.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Ultra-low-depth sequencing (ULDS) is a cost-effective strategy for large-scale genomic studies, but its utility hinges on accurate genotype imputation. This study evaluates three imputation tools -- STITCH, QUILT2, and GLIMPSE2 -- across varying sequencing depths and sample sizes, using the China Kadoorie Biobank (CKB) and The 1000 Genomes Project (1KGP) East Asian (EAS) reference panels. Critical performance divergences are demonstrated: Sample size sensitivity: STITCH's accuracy improved markedly with larger samples, whereas QUILT2 and GLIMPSE2 showed minimal dependence on sample size. Reference panel optimization: Population-specific CKB significantly enhanced accuracy for QUILT2 and GLIMPSE2 but had a negligible impact on STITCH, which relies on internal haplotype inference. Depth thresholds: All tools achieved robust accuracy at moderate sequencing depths (≥ 0.5x), but STITCH underperformed drastically at ultra-low depths (≤ 0.1x). GLIMPSE2 with CKB delivered the highest overall accuracy, while QUILT2 balanced precision and computational efficiency. For non-invasive prenatal testing (NIPT) data, GLIMPSE2+CKB maintained sufficient accuracy for downstream analyses. A decision framework is proposed, prioritizing population-matched panels and depth-adapted tools, offering actionable guidelines for optimizing ULDS-WGS in diverse research settings. These insights bridge methodological advancements with practical implementation, enabling cost-effective scaling of genomic studies without compromising data quality.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Ultra-low-depth sequencing (ULDS), defined as sequencing coverage below 1x, has gained traction due to its low cost, broad genome coverage, and compatibility with diverse sample types. It has already shown clinical value in applications such as non-invasive prenatal testing (NIPT)1, cancer monitoring2, and chromosomal copy number variation (CNV) detection3,4. Beyond clinical diagnostics, the decreasing cost of sequencing and rapid advances in bioinformatics have enabled ULDS to play a growing role in population genomics and complex trait research. By combining ULDS data with population-scale haplotype reference panels, genotype imputation enables the recovery of genome-wide variant information at the individual level. As a result, ULDS has emerged as a cost-effective alternative to traditional single-nucleotide polymorphism (SNP) arrays and high-depth whole-genome sequencing (WGS)5, particularly in large-scale studies such as genome-wide association studies (GWAS) and population structure analyses.

Previous research has demonstrated the feasibility of conducting various genetic studies using NIPT sequencing data, including variant calling, population history reconstruction, viral infection pattern inference, and GWAS6.

Despite these advantages, the extremely sparse nature of ULDS data presents unique challenges. At the variant level, many sites are entirely unobserved or represented by only a single allele per individual, leading to insufficient data quality for downstream analyses. Genotype imputation is therefore essential, leveraging haplotype structure from large reference panels (e.g., 1000 Genomes7 or population-specific resources) to statistically infer missing or uncertain genotypes. Previous work has shown that imputation from NIPT data can achieve high accuracy and retain robust statistical power in GWAS for identifying trait-associated variants8. Using the STITCH9 algorithm, NIPT data (mean depth ~0.15x) in a cohort of 20,900 Chinese pregnant women were successfully imputed, leading to the identification of pregnancy-associated loci. The imputed genotypes showed strong concordance with high-depth WGS data in GWAS results (Pearson R² > 0.8)10.

The success of ULDS-based analyses critically depends on imputation accuracy, which is influenced by sequencing depth, reference panel quality and population match, imputation algorithm performance, sample size, and allele frequency spectrum11. Among these, the choice of reference panel is a major determinant of imputation accuracy. Commonly used panels include globally representative resources such as the 1000 Genomes Project (1KGP)7, TOPMed12, and the Haplotype Reference Consortium (HRC)13, as well as increasingly available population- or region-specific panels such as Singapore 10,000 Genomes (SG10K)14, the China Kadoorie Biobank (CKB)15. Another key factor in imputation performance is the choice of algorithm. Several tools have been developed to accommodate the unique challenges of low-depth sequencing, significantly advancing the practical use of imputation in large-scale genetic research. While imputation methods such as Beagle (v5+)16, Minimac417, and IMPUTE511 are widely used for SNP array and medium-to-high-depth WGS data, they often perform suboptimally in ULDS settings. More recently, specialized tools have been developed to address these challenges. STITCH9 infers haplotypes directly from low-depth sequencing reads, making it particularly suitable for large homogeneous cohorts. QUILT218 employs a compressed haplotype library and localized likelihood model, enabling efficient imputation with massive reference panels and offering unique applications in prenatal genomics. GLIMPSE219, an extension of the original GLIMPSE framework, provides further improvements in both accuracy and computational efficiency.

Although these tools represent major advances, their relative performance under different experimental designs (e.g., sequencing depth, cohort size, and reference panel choice) has not been systematically evaluated, leaving researchers without clear guidance on selecting the most appropriate strategy. To bridge this gap, three widely used ULDS imputation tools -- STITCH, QUILT2, and GLIMPSE2 -- were systematically benchmarked under multiple sequencing depths and sample sizes. Their performance was evaluated using two East Asian reference panels highly relevant to Chinese populations. The findings indicate that ULDS imputation is generally reliable at sequencing depths ≥0.5x, while depths <0.1x require substantially larger cohorts to achieve acceptable accuracy. Reference panel selection should be tailored to the study context, with population-matched panels such as CKB improving imputation accuracy. Moreover, these approaches are directly applicable to ultra-low-depth data generated in large-scale population studies and NIPT. This study thus establishes a practical framework for tool selection in ULDS-based research, providing methodological guidance for future applications in population genetics and complex trait analyses.

Access restricted. Please log in or start a trial to view this content.

Protocol

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

All participants provided written informed consent prior to participation. The study involving high-depth WGS data was reviewed and approved by the BGI Institutional Review Board (BGI-IRB 23058-T2), and approval for the collection of human genetic resources was obtained from the Human Genetic Resources Administration of China ([2023] CJ0262). The study involving ULDS data from NIPT was approved by the Institutional Review Board of Wuhan Children's Hospital (2021R062) and the BGI Institutional Review Board (BGI-IRB 21088), with additional approval from the Human Genetic Resources Administration of China ([2021] CJ2002).

NOTE: This study included two types of WGS data. The first type consisted of high-depth WGS data (30xx) obtained from blood samples of 500 individuals recruited from a natural population cohort in Shenzhen. These data were used to construct a high-quality ground truth dataset and for subsequent down-sampling and accuracy evaluations. The second type comprised ULDS data derived from the NIPT of 10,000 pregnant women from the Wuhan area.

1. High-depth whole genome sequencing data

  1. Collect 500 peripheral blood samples (5 mL each) from a general population cohort after informed consent. Store samples in EDTA tubes and transport them at 2-8 °C.
  2. Centrifuge blood at 1,600 x g for 10 min at 4 °C to separate plasma and buffy coat. Carefully collect the buffy coat and store it at -80 °C until DNA extraction.
  3. Extract genomic DNA from buffy coat using a magnetic bead-based kit following the manufacturer's instructions.
  4. Quantify DNA concentration using a fluorometric assay and assess DNA integrity by agarose gel electrophoresis. Select samples with total DNA yield ≥1 µg, concentration ≥12.5 ng/µL, and fragment length >20 kb without visible degradation for library preparation.
  5. Shear 80-200 ng of high-quality genomic DNA to an average size of 350-400 bp by ultrasonication.
  6. Perform end-repair at 20 °C for 30 min, adapter ligation at 20 °C for 15 min, and circularization at 37 °C for 30 min to construct PCR-free libraries. Generate DNA nanoballs (DNBs) using rolling circle amplification (RCA). Sequence paired-end libraries (PE100, read length 100 bp) on a DNBSEQ platform to a target depth of ~30x (average 100 Gb per sample). Store raw sequencing reads in FASTQ format for downstream analysis.
    NOTE: Handle all human-derived samples under BSL-2 laboratory conditions. Avoid repeated freeze-thaw cycles to prevent DNA degradation. Dispose of blood-derived materials as biohazardous waste; dispose of chemical reagents following institutional hazardous waste guidelines.

2. Ultra-low depth NIPT data (~0.1x WGS)

  1. Collect 10,000 maternal blood samples (5 mL each) for routine non-invasive prenatal testing (NIPT). Use EDTA tubes and transport at 2-8 °C; process plasma within 8 h of collection.
  2. For stabilized circulating DNA tubes (K-tubes or G-tubes), transport at 6-35 °C using temperature-controlled carriers and process within 96 h following the manufacturer's standard operating procedures.
  3. Centrifuge blood at 1,600 x g for 10 min at 4 °C to separate plasma. Carefully collect the upper plasma layer without disturbing the buffy coat or cell pellet using a pipette and transfer it to a new tube. Centrifuge the recovered plasma again at 16,000 x g for 10 min at 4 °C to remove any residual cells or debris. Carefully transfer the clarified supernatant (cell-free plasma) to a fresh tube for DNA extraction.
  4. Extract circulating cell-free DNA (cfDNA) from plasma using a nucleic acid extraction kit. Perform end-repair at 20 °C for 30 min, adapter ligation at 20 °C for 15 min, and PCR amplification (12 cycles, 98 °C denaturation 10 s, 60 °C annealing 30 s, 72 °C extension 30 s).
  5. Purify PCR products and circularize libraries at 37 °C for 30 min. Generate DNBs through RCA. Sequence single-end libraries (SE35, read length 35 bp) on a BGISEQ-500 platform. Store raw sequencing data in FASTQ format.
    NOTE: Handle plasma samples as potentially infectious material under BSL-2 conditions. Minimize freeze-thaw cycles to reduce cfDNA degradation. Dispose of plasma waste and plastic consumables as biohazardous material.

3. Data preprocessing pipeline

  1. To systematically evaluate the performance of genotype imputation tools under varying sequencing depths, carry out a standardized preprocessing workflow on both the original high-depth WGS data (30x) and the ultra-low-depth NIPT data (<0.1x), including simulated downsampling, quality control, read alignment, duplicate removal, and base quality score recalibration (BQSR).
    NOTE: The steps from this point to imputation accuracy evaluation constitute the Main Protocol (Figure 1) of this study. The specific code can be found in Supplementary File 1.
  2. Downsampling
    1. Generate a series of downsampled datasets from the original 30x high-depth sequencing samples. Employ two strategies to realistically mimic the sequencing characteristics of NIPT data as described below.
    2. Random subsampling: Use seqtk v1.5 (https://github.com/lh3/seqtk) with a fixed random seed of 100 to create four levels of low-depth data (0.05x, 0.1x, 0.5x, and 1.0x).
    3. NIPT-like read structure simulation: Retain only the first read (R1) of each paired-end read and truncate all retained reads to 35 bp with seqtk trimfq -L 35, consistent with the typical single-end, short-read nature of ultra-low-depth NIPT sequencing.
  3. Quality control
    1. Process all raw FASTQ files with fastp v0.23.420. Use the following parameters: --qualified_quality_phred=5 (base quality threshold), --unqualified_percent_limit=50 (maximum percentage of low-quality bases allowed), --n_base_limit=10 (maximum N bases per read), and custom adapter removal with --adapter_sequence=AAGTCGGAGGCCAAGCGGTCTTAG
      GAAGACAA (R1) and --adapter_sequence_r2=AAGTCGGATCGTAGCC
      ATGTCGTTCTGTGAG
      CCAAGGAGTTG (R2).
    2. Disable Poly-G tail trimming (--disable_trim_poly_g), and generate reports in both JSON and HTML formats for each sample.
  4. Alignment and duplicate removal
    1. Align high-quality reads to the human reference genome GRCh38 (hg38)21 using BWA v0.7.16a-r118122.
    2. Perform the alignment with the aln algorithm (-e 10 -t 4 -i 5 -q 0), followed by samse for single-end alignment with read group information.
    3. Convert the resulting SAM files to BAM, sorted (samtools sort -@ 8), and remove duplicates using SAMtools v1.323 (samtools rmdup). Index all BAM files.
  5. Base quality score recalibration (BQSR)
    1. Perform BQSR using GATK v4.0.4.024. Train the recalibration model on three high-confidence variant datasets: dbSNP build 14625, Mills and 1000G gold standard indels21, and the GATK resource bundle24 known indels file for GRCh38. The bundle files used refer to the GATK official example (https://github.com/gatk-workflows/gatk4-data-processing/blob/master/processing-for-variant-discovery-gatk4.hg38.wgs.inputs.json). In total, download three files and their corresponding index files.
    2. Run BaseRecalibrator followed by ApplyBQSR to generate recalibrated BAM files. Index all BAMs using SAMtools v1.3.
      NOTE: All simulated datasets underwent identical preprocessing steps-downsampling, quality control, alignment, duplicate removal, and BQSR-to ensure consistency and comparability in subsequent imputation performance evaluations.

4. Genotype imputation

  1. Data preparation
    1. Imputation dataset setup: Construct multiple evaluation datasets to systematically benchmark genotype imputation tools across different depths and sample sizes as described below. A total of nine combinations are formed based on the different sample sizes and sequencing depths mentioned above. The input files consist of the sequencing data BAM file lists (bamlist.txt) for the above-mentioned nine subsets after quality control, stored in corresponding bamlist.txt files. Other input files include the human reference genome (GRCh3821) and the genetic map from the 1000 Genomes Project7.
      1. Downsampled high-depth WGS data: Randomly select two subsets (200 and 500 samples) from 500 individuals sequenced at 30x depth. Downsample each subset to four depths (1x, 0.5x, 0.1x, and 0.05x) to generate eight experimental conditions.
      2. NIPT-based ULDS dataset: Combine 10,000 ultra-low-depth NIPT samples (average depth is 0.102x, Figure 2) with 50 high-depth samples downsampled to 0.1x.
    2. Analysis region specification: Restrict all analyses to the chromosome 1 region chr1:150,500,000-160,500,000 (10 Mb), with a 500 kb buffer for imputation, to ensure direct comparability across tools.
    3. Reference panel selection: Use two reference panels (Table 1): the CKB panel, constructed from Chinese population data, and the 1KGP-EAS panel, derived from the East Asian subset of the 1000 Genomes Project.
      NOTE: The CKB reference panel15 was constructed using high-depth (~15x) whole-genome sequencing data from 9,964 Chinese adults in the China Kadoorie Biobank, a large prospective cohort study. These samples are derived from a natural population with minimal phenotypic bias, homogeneous Han Chinese ancestry, and consistent population structure, making them particularly well-suited for genotype imputation in Chinese cohorts. Yu et al. 15 demonstrated that, in a real-phenotype GWAS for height, imputation using the CKB panel tripled the number of SNPs detected and doubled the number of genome-wide significant variants. The 1000 Genomes Project (1KGP)7,26, the most widely used genomic reference, includes 585 individuals in its East Asian (EAS) Phase 3 subset. This subset covers five East Asian populations, which have a sequencing depth of approximately 30x, including Han Chinese in Beijing (CHB), Southern Han Chinese (CHS), Chinese Dai in Xishuangbanna (CDX), Kinh in Ho Chi Minh City, Vietnam (KHV), and Japanese in Tokyo (JPT).
  2. Imputation tools
    1. Evaluate three imputation algorithms, chosen for their distinct modeling strategies and applicability to ultra-low-depth sequencing (ULDS) data.
      1. STITCH: STITCH (v1.6.6) is a reference-free haplotype-based imputation algorithm that can optionally incorporate external reference haplotypes. Include BAM lists, human reference genome (GRCh38) as input files. Prepare reference panel files (hap/legend/pos) when performing reference-based imputation. Include the following key parameters: method=diploid, buffer=500 kb, K=10 ancestral haplotypes, and nGen=4x sample size/K (as recommended in STITCH documentation). Generate output files containing per-SNP genotype dosages for all individuals.
        NOTE: As per the STITCH official documentation, K is the number of ancestral haplotypes in the model. A larger K improves imputation accuracy for larger samples and higher coverages, but it also increases computation time, and accuracy may decrease with lower coverage.
      2. QUILT2: QUILT2 applies a Bayesian reference-guided approach optimized for ULDS data. Perform reference panel preparation using the provided prepare_reference script, specifying the genetic map and region coordinates. Run imputation diploid mode with the same buffer size (500 kb) and nGen setting as STITCH to ensure comparability.
      3. GLIMPSE2: GLIMPSE2 is an HMM-based reference-driven imputation tool, designed for large-scale and very low-depth sequencing datasets. Execute imputation with GLIMPSE2_phase_static, specifying input BAM list, human reference VCF panel, genetic map, input region = chr1:150,000,000-161,000,000, output region = chr1:150,500,000-160,500,000 (to maintain a 500 kb buffer). The BAM list file entered here needs to contain two columns: one is the BAM path, and the second column is the sample name. If the second column is not entered, each BAM file name will be used as the sample name in the output VCF file.

5. Imputation accuracy evaluation

  1. Truth set definition
    1. Select 50 individuals sequenced at 30x depth as the ground-truth dataset. Include these samples in the experimental downsampling conditions to ensure comparability.
    2. Perform variant calling for the truth set using the following pipeline: SOAPnuke27 for QC, BWA for alignment, Picard for duplicate marking, GATK v4.0.4.0 BQSR and HaplotypeCaller for variant calling, DPGT (https://github.com/BGI-flexlab/DPGT) for joint calling, and BCFtools v1.1123 for variant quality filtering.
    3. Retain only high-confidence PASS variants to generate a benchmark VCF file. Restrict the evaluation to chr1:150,500,000-160,500,000, consistent with the imputed datasets.
  2. Data harmonization and filtering
    1. Process imputed VCF files from all tools using PLINK2.028. Extract dosage data and convert to pgen format.
    2. Apply SNP-level quality control with the following filters: Minor allele frequency (MAF) ≥ 0.05 (--maf 0.05), Hardy-Weinberg equilibrium (HWE) p-value ≥ 1e-6 (--hwe 1e-6), Biallelic SNPs only (--max-alleles 2).
    3. Export variants passing QC to traw format for downstream comparison.
  3. Accuracy metrics
    1. Compare the imputed dosages with the ground-truth dosages for each SNP. Compute Pearson correlation coefficients (R) on a SNP-by-SNP basis, retain only sites common to both datasets, and calculate the mean of squared correlation coefficients (R²) across all evaluated SNPs to quantify overall imputation accuracy for each condition.
    2. Use this accuracy metric to capture the concordance of genotype dosage estimates between imputed and true genotypes.

Access restricted. Please log in or start a trial to view this content.

Results

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Impact of sample size on imputation accuracy
Increasing the sample size from N = 200 to N = 500 improved the imputation accuracy of STITCH, particularly under low-coverage conditions. For example, with the CKB reference panel at 1x coverage, STITCH achieved an R2> of 0.916 (N=500) compared to 0.882 (N=200), representing a 3.4% increase (Figure 3; Supplementary File 2). Similarly, at 0.5x coverage, its accuracy rose from 0.8...

Access restricted. Please log in or start a trial to view this content.

Discussion

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study systematically assessed the performance of three widely used genotype imputation tools for ULDS, with high-depth WGS serving as the gold standard. A key methodological strength lies in the adoption of a unified preprocessing pipeline -- encompassing alignment, quality control, and base quality score recalibration -- that minimizes batch effects and ensures comparability across tools and conditions. By downsampling deeply sequenced samples, ultra-low-depth data were simulated under controlled settings, thereby ...

Access restricted. Please log in or start a trial to view this content.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors declare no competing interests.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study was supported by Shenzhen Medical Research Fund (B2404004), National Key Research and Development Program of China (2023YFC2605400, 2022YFC2502402), Shenzhen Science and Technology Program (SYSPG20241211173852024), Open Research Project in State Key Laboratory of Vascular Homeostasis and Remodeling (Peking University) (2025-SKLVHR-013), and Key-Area Research and Development Program of Guangdong Province (2023B0303040001).

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
Data
10,000 NIPT low-depth samplesThis paperUltra-low-depth whole genome sequencing data used for imputation benchmark.
500 high-depth WGS samplesThis paper30× high-depth WGS used as gold standard/truth set.
Reference Panel
1KGP-EAS reference panel1000 Genomes Project (East Asia)Subset of 1KGP for East Asian ancestry-specific imputation.
CKB reference panelChina Kadoorie BiobankCustom population-specific panel for genotype imputation.
Software and algorithms
BCFtools v1.11GitHub (samtools/bcftools)Used for merging and sorting chromosome-level results, and filtering variants.
BQSR of the GATK 4.0.4.0 toolsetBroad InstituteUsed for base quality score recalibration (BQSR).
BWA-MEM .7.16a-r1181Heng Li / GitHubFor aligning raw reads to GRCh38.
DPGT (Distributed Population Genetics Tool)BGIA distributed population genetics analysis tool which enabled joint calling on millions of WGS samples. Available at [GitHub - BGI-flexlab/DPGT](https://github.com/BGI-flexlab/DPGT)
fastp.0.23.4Open-source (Chen et al., 2018)For quality control and adapter trimming.
GLIMPSE2University of OxfordFast genotype phasing and imputation for low-coverage WGS
Original Code for analysesThis paperSupplementary File 1 Original Code for analyses
Picard toolkitBroad InstituteUsed for marking duplicates and file format conversion.
Plink 2.0C. Chang, S. Purcell / Broad InstituteFor genotype format conversion and association analysis.
Python 3.8Python Software FoundationUsed for scripting, automation, and data analysis.
QUILT2Oxford Big Data InstituteHMM-based imputation using external reference panels
R 4.1.3The R FoundationUsed for running STITCH, QUILT2 and plotting/statistics.
SAMtools v1.3GitHub (samtools/samtools)For manipulating SAM/BAM files.
Seqtk-1.5GitHub (lh3/seqtk)Toolkit for processing sequences in FASTA/Q formats. Available at [GitHub - lh3/seqtk](https://github.com/lh3/seqtk)
SOAPnukeBGIFor NGS data quality control and filtering.
STITCH v1.6.6University of OxfordImputation tool optimized for ultra-low coverage sequencing
tabixGitHub (samtools/tabix)Used for indexing and querying bgzipped VCF files.
Other Materials
GATK bundle filesGATKAvailable at [https://github.com/gatk-workflows/gatk4-data-processing/blob/master/processing-for-variant-discovery-gatk4.hg38.wgs.inputs.json]
Genetic Map for 1000G (GRCh38)Oxford / 1000 Genomes ProjectRequired for phasing/imputation tools
GRCh38Genome Reference ConsortiumUsed for read alignment and variant calling

References

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,
  1. Zhang, H., et al. Non-invasive prenatal testing for trisomies 21, 18 and 13: clinical experience from 146,958 pregnancies. Ultrasound Obstet Gynecol. 45 (5), 530-538 (2015).
  2. Heitzer, E., et al. Tumor-associated copy number changes in the circulation of patients with prostate cancer identified through whole-genome sequencing. Genome Med. 5 (4), 30(2013).
  3. Hyblova, M., et al. Validation of Copy Number Variants Detection from Pregnant Plasma Using Low-Pass Whole-Genome Sequencing in Noninvasive Prenatal Testing-Like Settings. Diagnostics (Basel). 10 (8), (2020).
  4. Kucharik, M., Budis, J., Hyblova, M., Minarik, G., Szemes, T. Copy Number Variant Detection with Low-Coverage Whole-Genome Sequencing Represents a Viable Alternative to the Conventional Array-CGH. Diagnostics (Basel). 11 (4), (2021).
  5. Li, J. H., Mazur, C. A., Berisa, T., Pickrell, J. K. Low-pass sequencing increases the power of GWAS and decreases measurement error of polygenic risk scores compared to genotyping arrays. Genome Res. 31 (4), 529-537 (2021).
  6. Liu, S., et al. Genomic Analyses from Non-invasive Prenatal Testing Reveal Genetic Associations, Patterns of Viral Infections, and Chinese Population History. Cell. 175 (2), 347-359.e14 (2018).
  7. The 1000 Genomes Project Consortium. A global reference for human genetic variation. Nature. 526 (7571), 68-74 (2015).
  8. Liu, S., et al. Utilizing non-invasive prenatal test sequencing data for human genetic investigation. Cell Genom. 4 (10), 100669(2024).
  9. Davies, R. W., Flint, J., Myers, S., Mott, R. Rapid genotype imputation from sequence without reference panels. Nat Genet. 48 (8), 965-969 (2016).
  10. Xiao, H., et al. Genetic analyses of 104 phenotypes in 20,900 Chinese pregnant women reveal pregnancy-specific discoveries. Cell Genom. 4 (10), 100633(2024).
  11. Rubinacci, S., Ribeiro, D. M., Hofmeister, R. J., Delaneau, O. Efficient phasing and imputation of low-coverage sequencing data using large reference panels. Nat Genet. 53 (1), 120-126 (2021).
  12. Taliun, D., et al. Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program. Nature. 590 (7845), 290-299 (2021).
  13. McCarthy, S., et al. A reference panel of 64,976 haplotypes for genotype imputation. Nat Genet. 48 (10), 1279-1283 (2016).
  14. Wu, D., et al. Large-Scale Whole-Genome Sequencing of Three Diverse Asian Populations in Singapore. Cell. 179 (3), 736-749.e15 (2019).
  15. Yu, C., et al. A high-resolution haplotype-resolved Reference panel constructed from the China Kadoorie Biobank Study. Nucleic Acids Res. 51 (21), 11770-11782 (2023).
  16. Browning, B. L., Zhou, Y., Browning, S. R. A One-Penny Imputed Genome from Next-Generation Reference Panels. Am J Hum Genet. 103 (3), 338-348 (2018).
  17. Das, S., et al. Next-generation genotype imputation service and methods. Nat Genet. 48 (10), 1284-1287 (2016).
  18. Li, Z., Albrechtsen, A., Davies, R. W. Rapid and accurate genotype imputation from low coverage short read, long read, and cell free DNA sequence. bioRxiv. , (2024).
  19. Rubinacci, S., Ribeiro, D. M., Hofmeister, R. J., Delaneau, O. Efficient phasing and imputation of low-coverage sequencing data using large reference panels. Nat Genet. 53 (1), 120-126 (2021).
  20. Chen, S. Ultrafast one-pass FASTQ data preprocessing, quality control, and deduplication using fastp. Imeta. 2 (2), e107(2023).
  21. Zheng-Bradley, X., et al. Alignment of 1000 Genomes Project reads to reference assembly GRCh38. Gigascience. 6 (7), 1-8 (2017).
  22. Li, H., Durbin, R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics. 25 (14), 1754-1760 (2009).
  23. Danecek, P., et al. Twelve years of SAMtools and BCFtools. Gigascience. 10 (2), 8(2021).
  24. Van der Auwera, G. A., et al. From FastQ data to high confidence variant calls: the Genome Analysis Toolkit best practices pipeline. Curr Protoc Bioinfo. 43 (1110), 11.10.1-11.10.33 (2013).
  25. Sherry, S. T., et al. dbSNP: the NCBI database of genetic variation. Nucleic Acids Res. 29 (1), 308-311 (2001).
  26. Byrska-Bishop, M., et al. High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios. Cell. 185 (18), 3426-3440 (2022).
  27. Chen, Y., et al. SOAPnuke: a MapReduce acceleration-supported software for integrated quality control and preprocessing of high-throughput sequencing data. Gigascience. 7 (1), 1-6 (2018).
  28. Chang, C. C., et al. Second-generation PLINK: rising to the challenge of larger and richer datasets. Gigascience. 4 (7), 8(2015).
  29. Marchini, J., Howie, B. Genotype imputation for genome-wide association studies. Nat Rev Genet. 11 (7), 2796(2010).
  30. Zeng, J., et al. Protocol for genetic analysis of population-scale ultra-low-depth sequencing data. STAR Protoc. 6 (1), 579(2025).
  31. Naito, T., Okada, Y. Genotype imputation methods for whole and complex genomic regions utilizing deep learning technology. J Hum Genet. 69 (10), 481-486 (2024).

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

Genotype ImputationUltra Low Depth SequencingWhole Genome SequencingImputation ToolsReference Panel OptimizationPopulation Specific PanelsSample Size SensitivityNon Invasive Prenatal TestingSequencing Depth ThresholdsComputational Efficiency

Related Articles