$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Ultra-low-depth sequencing (ULDS), defined as sequencing coverage below 1x, has gained traction due to its low cost, broad genome coverage, and compatibility with diverse sample types. It has already shown clinical value in applications such as non-invasive prenatal testing (NIPT)1, cancer monitoring2, and chromosomal copy number variation (CNV) detection3,4. Beyond clinical diagnostics, the decreasing cost of sequencing and rapid advances in bioinformatics have enabled ULDS to play a growing role in population genomics and complex trait research. By combining ULDS data with population-scale haplotype reference panels, genotype imputation enables the recovery of genome-wide variant information at the individual level. As a result, ULDS has emerged as a cost-effective alternative to traditional single-nucleotide polymorphism (SNP) arrays and high-depth whole-genome sequencing (WGS)5, particularly in large-scale studies such as genome-wide association studies (GWAS) and population structure analyses.
Previous research has demonstrated the feasibility of conducting various genetic studies using NIPT sequencing data, including variant calling, population history reconstruction, viral infection pattern inference, and GWAS6.
Despite these advantages, the extremely sparse nature of ULDS data presents unique challenges. At the variant level, many sites are entirely unobserved or represented by only a single allele per individual, leading to insufficient data quality for downstream analyses. Genotype imputation is therefore essential, leveraging haplotype structure from large reference panels (e.g., 1000 Genomes7 or population-specific resources) to statistically infer missing or uncertain genotypes. Previous work has shown that imputation from NIPT data can achieve high accuracy and retain robust statistical power in GWAS for identifying trait-associated variants8. Using the STITCH9 algorithm, NIPT data (mean depth ~0.15x) in a cohort of 20,900 Chinese pregnant women were successfully imputed, leading to the identification of pregnancy-associated loci. The imputed genotypes showed strong concordance with high-depth WGS data in GWAS results (Pearson R² > 0.8)10.
The success of ULDS-based analyses critically depends on imputation accuracy, which is influenced by sequencing depth, reference panel quality and population match, imputation algorithm performance, sample size, and allele frequency spectrum11. Among these, the choice of reference panel is a major determinant of imputation accuracy. Commonly used panels include globally representative resources such as the 1000 Genomes Project (1KGP)7, TOPMed12, and the Haplotype Reference Consortium (HRC)13, as well as increasingly available population- or region-specific panels such as Singapore 10,000 Genomes (SG10K)14, the China Kadoorie Biobank (CKB)15. Another key factor in imputation performance is the choice of algorithm. Several tools have been developed to accommodate the unique challenges of low-depth sequencing, significantly advancing the practical use of imputation in large-scale genetic research. While imputation methods such as Beagle (v5+)16, Minimac417, and IMPUTE511 are widely used for SNP array and medium-to-high-depth WGS data, they often perform suboptimally in ULDS settings. More recently, specialized tools have been developed to address these challenges. STITCH9 infers haplotypes directly from low-depth sequencing reads, making it particularly suitable for large homogeneous cohorts. QUILT218 employs a compressed haplotype library and localized likelihood model, enabling efficient imputation with massive reference panels and offering unique applications in prenatal genomics. GLIMPSE219, an extension of the original GLIMPSE framework, provides further improvements in both accuracy and computational efficiency.
Although these tools represent major advances, their relative performance under different experimental designs (e.g., sequencing depth, cohort size, and reference panel choice) has not been systematically evaluated, leaving researchers without clear guidance on selecting the most appropriate strategy. To bridge this gap, three widely used ULDS imputation tools -- STITCH, QUILT2, and GLIMPSE2 -- were systematically benchmarked under multiple sequencing depths and sample sizes. Their performance was evaluated using two East Asian reference panels highly relevant to Chinese populations. The findings indicate that ULDS imputation is generally reliable at sequencing depths ≥0.5x, while depths <0.1x require substantially larger cohorts to achieve acceptable accuracy. Reference panel selection should be tailored to the study context, with population-matched panels such as CKB improving imputation accuracy. Moreover, these approaches are directly applicable to ultra-low-depth data generated in large-scale population studies and NIPT. This study thus establishes a practical framework for tool selection in ULDS-based research, providing methodological guidance for future applications in population genetics and complex trait analyses.