Impact of sample size on imputation accuracy
Increasing the sample size from N = 200 to N = 500 improved the imputation accuracy of STITCH, particularly under low-coverage conditions. For example, with the CKB reference panel at 1x coverage, STITCH achieved an R2> of 0.916 (N=500) compared to 0.882 (N=200), representing a 3.4% increase (Figure 3; Supplementary File 2). Similarly, at 0.5x coverage, its accuracy rose from 0.800 to 0.868 (ΔR2> = 8.5%). In contrast, QUILT2 and GLIMPSE2 showed minimal sensitivity to sample size variations, with R2> fluctuations of less than 0.5% across all tested conditions. For instance, QUILT2 maintained stable performance under the CKB panel at 1x coverage (R2> = 0.970 for N = 200 vs. 0.971 for N = 500). This indicates that STITCH benefits more from larger sample sizes, consistent with its hidden Markov model (HMM)-based haplotype inference framework.
Reference panel compatibility
The choice of reference panel affected the imputation accuracy of QUILT2 and GLIMPSE2, while having little impact on STITCH (Figure 3; Supplementary File 2). When using the Chinese population-specific CKB panel, both QUILT2 and GLIMPSE2 achieved higher accuracy than with the general-purpose EAS panel. For example, at 1x coverage, the R2> of GLIMPSE2 increased from 0.956 with EAS to 0.974 with CKB (ΔR2> = 1.8%), while QUILT2 rose from 0.950 to 0.971 (ΔR2> = 2.1%). In contrast, STITCH showed almost no difference between the two panels (R2> = 0.916 versus 0.915), consistent with its independence from external reference panels. These results highlight the practical advantage of population-specific reference panels in precision medicine research.
Sequencing depth and accuracy trade-offs
The imputation accuracy of all tools increased with higher sequencing depth, but differences between methods became more evident under ultra-low coverage conditions (≤ 0.1x) (Figure 3; Supplementary File 2). Using the CKB panel as an example, GLIMPSE2 achieved an R2> of 0.676 at 0.05x coverage, which increased to 0.974 at 1x. Similarly, QUILT2 rose from 0.666 to 0.971 over the same range. In contrast, STITCH performed poorly at 0.05x coverage, with an R2> of 0.255, indicating reduced robustness under extremely low coverage. At 0.1x coverage, QUILT2 and GLIMPSE2 reached R2> values above 0.817 with the CKB panel, whereas STITCH reached only 0.532 (N = 500).
SNP count considerations and impact on tool selection
STITCH produced fewer total SNPs than QUILT2 and GLIMPSE2, but after quality control, the differences in effective SNP numbers among the tools were reduced. The total number of imputed SNPs varied across tools. For example, at a sample size of N = 500 and coverage of 1x, STITCH generated 21,829 SNPs with the CKB panel and 24,546 with the EAS panel, compared with 398,026 (CKB) and 205,711 (EAS) for QUILT2, and 467,754 (CKB) and 239,495 (EAS) for GLIMPSE2 under the same conditions. After applying standard quality control procedures, MAF and HWE filtering, the number of SNPs retained for downstream analysis became comparable across all tools (Figure 4; Supplementary File 2). This suggests that although the raw SNP counts differ across methods, the number of usable variants after QC is similar, and therefore total imputed SNP count should not be considered a decisive factor when selecting an imputation tool.
Computational resource consumption and tool selection recommendations
In this study, core-hours -- defined as the product of the number of CPU cores and runtime -- were used as a key metric to quantify the computational burden of genotype imputation tools. Under the same sample size and sequencing depth, imputation using the CKB reference panel consistently required more core-hours than using the EAS panel, suggesting that larger or more complex reference panels increase computational demand. Additionally, core-hour usage increased with both sample size and sequencing depth, highlighting the need for careful resource planning in large-scale studies (Figure 5; Supplementary File 3). Among the three tools evaluated, QUILT2 consumed substantially more core-hours than both STITCH and GLIMPSE2, especially at the largest sample size (N = 10,050), where QUILT2 required 361-496 core-hours per imputation run, compared to approximately 20-22 for STITCH and 87-145 for GLIMPSE2. Based on these observations, GLIMPSE2 is recommended for small- to medium-sized datasets, particularly when imputation accuracy is a primary concern, but computational resources are constrained. Conversely, STITCH should be considered the tool of choice in studies that lack high-quality reference panels or aim to handle very large cohorts, as it offers a favorable trade-off between accuracy and efficiency. Researchers are therefore advised to balance accuracy, reference panel availability, and computational costs when selecting imputation tools for ultra-low-depth WGS studies.

Figure 1: Main steps of the comprehensive evaluation of genotype imputation tools. This flowchart illustrates the three main steps of genotype imputation accuracy evaluation: Data Preprocessing Pipeline, Genotype Imputation Workflow, and Imputation Accuracy Evaluation. Please click here to view a larger version of this figure.

Figure 2: Sequencing depth distribution of 10,000 ultra-low-depth NIPT samples. Histogram showing the coverage depth distribution of 10,000 non-invasive prenatal testing (NIPT) samples sequenced at ultra-low depth. The red dashed line indicates the mean coverage depth (0.102x), and the green dashed line represents the median coverage depth (0.104x). The x-axis ranges from 0.05x to 0.35x, with the majority of samples concentrated around 0.10x. Please click here to view a larger version of this figure.

Figure 3: Evaluation of imputation accuracy for three genotype imputation tools under varying conditions. (A) Accuracy across varying coverage depths (0.05x, 0.1x, 0.5x, 1x) using two reference panels (CKB versus EAS), with sample size fixed at 500. (B) Accuracy across varying coverage depths under two sample sizes (200 versus 500), using the EAS reference panel. (C) Accuracy on 0.1x coverage depth data under different sample sizes (200, 500, and 10,050). Please click here to view a larger version of this figure.

Figure 4: Comparison of SNP numbers from different genotype imputation tools. A sample size of n = 500 and coverage depth of 1x: (A) The total number of SNPs predicted and the number of SNPs remaining after quality control (MAF > 0.05 and HWE filtering) by three genotype imputation tools (STITCH, QUILT2, and GLIMPSE2) in the EAS and CKB populations. (B) The number of SNPs after quality control and the number of effective SNPs used for calculating imputation accuracy. Please click here to view a larger version of this figure.

Figure 5: Comparison of core-hour consumption across tools under varying coverage depths, sample sizes, and reference panels. (A) Core-hour consumption changes with increasing coverage depth under a fixed sample size (N = 500), comparing three tools (STITCH, GLIMPSE2, and QUILT2) using the EAS and CKB reference panels. (B) Compares core-hour usage across different sample sizes (200 vs. 500) as coverage depth increases, under the EAS reference panel. (C) The impact of sample size (200, 500, 10,050) on core-hour consumption at a fixed coverage depth of 0.1x, under both EAS and CKB reference panels. Please click here to view a larger version of this figure.
| Reference Panel | Sample size | Sequencing depth | Ancestries |
| CKB | 9,964 | ~15X | Chinese |
| 1KG-EAS | 585 | ~30X | Five East Asian populations |
Table 1: Information on reference panels. The table provides key details about the reference panels used for genotype imputation, including the name of the reference panel, sample size, sequencing depth, and ancestries.
| Tool | Strengths | Limitations | Recommended Scenarios |
| STITCH | Reference-free; low computational burden; scalable to very large cohorts | Lower accuracy, especially at ultra-low depth | Large-scale cohorts (N ≥ 10,000); populations without appropriate reference panels |
| QUILT2 | High accuracy; robust across depths; supports maternal–fetal separation | Higher computational cost; under active development | Ultra-low-depth NIPT data; studies requiring maternal–fetal genome separation |
| GLIMPSE2 | Highest accuracy (R2> up to 0.974 with CKB at 1×); efficient resource usage | Requires reference panel; slightly less robust at extremely low depths | Small- to medium-sized datasets; studies prioritizing accuracy with available reference |
Table 2: Recommended application scenarios for STITCH, QUILT2, and GLIMPSE2. Recommended application scenarios for STITCH, QUILT2, and GLIMPSE2. Recommendations are derived from comparative evaluations under varying sequencing depths, sample sizes, and reference panel conditions, aiming to guide tool selection in ultra-low-depth WGS studies.
Supplementary File 1: Original code for analyses. Please click here to download this File.
Supplementary File 2: Imputation accuracy and SNPs number. Please click here to download this File.
Supplementary File 3: Core-hour consumption of imputation. Please click here to download this File.