Anchor-step benchmarking separated tissue-specific from pan-tissue gene effects
The curated GSE117827 anchor matrix contained 27,685 HGNC-approved genes. The primary nasal contrast included 15 symptomatic infected and 6 virus-negative control samples; the blood contrast included 13 symptomatic infected and 6 controls (Table 1). Because this small cohort cannot support stable single-gene discovery, it was used as a paired-compartment anchor. Stability was evaluated at the module level by bootstrap resampling, random-gene null testing, external validation, and module-size sensitivity analyses.
| Source | Group | Condition | Primary contrast | n |
| Blood | RSV | Infected | Yes | 4 |
| Blood | Asymptomatic picornavirus | Secondary | No | 5 |
| Blood | Symptomatic picornavirus | Infected | Yes | 9 |
| Blood | Virus-negative control | Control | Yes | 6 |
| Nasal | RSV | Infected | Yes | 6 |
| Nasal | Asymptomatic picornavirus | Secondary | No | 5 |
| Nasal | Symptomatic picornavirus | Infected | Yes | 9 |
| Nasal | Virus-negative control | Control | Yes | 6 |
Table 1: GSE117827 anchor design after primary contrast curation. Sample distribution across blood and nasal compartments in the paired anchor dataset after primary contrast curation. Symptomatic respiratory syncytial virus (RSV) and symptomatic picornavirus samples were classified as infected and included in the primary contrast, whereas virus-negative controls were classified as controls and included in the primary contrast. Asymptomatic picornavirus samples were designated as secondary and excluded from module construction; they were used only in the exploratory symptom-gradient analysis. Two RSV cases lacked blood specimens, resulting in four blood and six nasal RSV samples.
Across the 27,685 shared genes, nasal and blood Hedges g estimates were nearly uncorrelated (Pearson r = 0.015; Figure 1). This result arose within one study and is less exposed to between-study differences than a pooled cross-cohort comparison. The dense central cloud shows that most genes did not move similarly in both compartments. The highlighted overlap genes were exceptions at the top of both rank lists. This pattern supports separate nasal and blood module construction.

Figure 1. Nasal and blood gene-level effect sizes in the paired GSE117827 anchor cohort. Hedges g values compare symptomatic infection with virus-negative controls for 27,685 genes. Nasal estimates (15 infected, 6 controls) are shown on the x-axis, and blood estimates (13 infected, 6 controls) are shown on the y-axis. Hexagonal-bin shading indicates the number of genes per bin. Red points indicate the six genes shared between the top-50 nasal and blood modules. Pearson r = 0.015. Please click here to view a larger version of this figure.
Module construction produced a small interferon-centered overlap
The top 50 nasal and blood modules shared six genes: ISG15, ACRBP, IFIT1, RSAD2, CCRL2, and XAF1 (Jaccard index = 0.064; Table 2; Figure 2). ISG15, IFIT1, RSAD2, and XAF1 are compatible with interferon-linked antiviral biology32,33,34. CCRL2 is better interpreted in the context of inflammatory leukocyte migration35. ACRBP has no established antiviral role. All six genes are reported for transparency, but none is claimed as a validated cross-tissue biomarker. Their single-gene FDR values were not significant in the small anchor cohort. The main inference therefore rests on locked module-level validation, not on the overlap list.
| Gene | Nasal Hedges g | Nasal FDR | Blood Hedges g | Blood FDR | Interpretation |
| ISG15 | 2.215 | 0.213 | 1.369 | 0.442 | Interferon-stimulated antiviral gene; exploratory overlap |
| ACRBP | 1.69 | 0.205 | 1.748 | 0.429 | No established antiviral role; retained for transparency |
| IFIT1 | 1.875 | 0.213 | 1.35 | 0.442 | Interferon-stimulated antiviral gene; exploratory overlap |
| RSAD2 | 1.663 | 0.213 | 1.532 | 0.442 | Interferon-stimulated antiviral gene; exploratory overlap |
| CCRL2 | 1.396 | 0.217 | 1.782 | 0.442 | Inflammatory leukocyte-migration context; not virus-specific |
| XAF1 | 1.603 | 0.226 | 1.47 | 0.442 | Interferon-linked apoptosis factor; exploratory overlap |
Table 2: Complete exploratory overlap between the top-ranked nasal and blood modules. All six shared genes are reported with nasal and blood Hedges g and gene-wise FDR values. All six genes are treated as exploratory rank-overlap candidates because the gene-wise FDR values were not significant in the small anchor cohort. ACRBP is retained for transparency despite lacking an established antiviral role. None of the six genes is presented as a validated universal biomarker; the primary evidence is based on module-level external validation.

Figure 2. Effect sizes of the six exploratory nasal-blood overlap genes. Nasal and blood Hedges g estimates are shown for ACRBP, CCRL2, IFIT1, ISG15, RSAD2, and XAF1 in GSE117827. Positive values indicate higher expression in symptomatic infection than in virus-negative controls. The complete overlap is shown for transparency; single-gene FDR values were not significant, and these genes are not presented as validated universal biomarkers. Please click here to view a larger version of this figure.
Functional enrichment provided a biological check on module content
Both modules were enriched for interferon and antiviral pathways, although the gene composition and enrichment strength differed (Figure 3). The eight most significant terms per module are displayed. For the nasal module, Hallmark interferon-gamma response (adjusted P = 2.23 × 10⁻24), Hallmark interferon-alpha response (adjusted P = 1.40 × 10⁻22), Reactome interferon-alpha/beta signalling (adjusted P = 3.64 × 10⁻19), and GO defense response to virus (adjusted P = 5.47 × 10⁻15) ranked among the leading terms. The blood module showed the same broad biology at lower enrichment strength: interferon-alpha response (adjusted P = 3.87 × 10⁻6), Reactome interferon-alpha/beta signaling (adjusted P = 5.70 × 10⁻6), interferon-gamma response (adjusted P = 8.89 × 10⁻6), and defense response to virus (adjusted P = 1.07 × 10⁻4). These results support biological coherence without implying identical compartmental gene rankings.

Figure 3. Selected enrichment terms for the nasal and blood modules. Over-representation analysis was performed using Enrichr with Hallmark 2020, Reactome 2022, and GO Biological Process 2023. The eight terms with the smallest Benjamini–Hochberg-adjusted P values for each module are shown. Bar length represents −log10(adjusted P value). The left and right panels show the nasal and blood modules, respectively. Terms are displayed in sentence case. Enrichment analysis did not alter module membership. Please click here to view a larger version of this figure.
Locked modules were tested in tissue-matched external validation
The locked nasal module reached AUROCs of 0.749 in GSE41374, 0.693 in GSE152075, and 0.609 in GSE156063 (Table 3; Figure 4). The locked blood module reached AUROCs of 0.832 in GSE171110, 0.924 in the GSE38900 GPL10558 unit, and 0.870 in the GPL6884 unit. Internal anchor performance is omitted because it is optimistically biased by construction. External AUROCs are treated as portability summaries. They do not establish clinical sensitivity, specificity, or readiness for diagnosis.
| Cohort | Sample/virus | Module | n positive | n negative | Genes represented | AUROC | Average precision | Welch-test FDR |
| GSE152075 | SARS-CoV-2 upper airway | Nasal | 430 | 54 | 50 | 0.693 | 0.935 | 2.20 × 10⁻⁴ |
| GSE156063 | SARS-CoV-2 upper airway | Nasal | 93 | 100 | 49 | 0.609 | 0.551 | 1.38 × 10⁻² |
| GSE171110 | SARS-CoV-2 whole blood | Blood | 44 | 10 | 50 | 0.832 | 0.959 | 7.94 × 10⁻⁴ |
| GSE38900-GPL10558 | RSV whole blood | Blood | 28 | 8 | 50 | 0.924 | 0.979 | 7.94 × 10⁻⁴ |
| GSE38900-GPL6884 | RSV whole blood | Blood | 107 | 31 | 49 | 0.87 | 0.962 | 3.47 × 10⁻¹⁵ |
| GSE41374 | RSV nasal wash | Nasal | 76 | 10 | 50 | 0.749 | 0.951 | 2.63 × 10⁻² |
Table 3: External tissue-matched module-score validation. The locked nasal module was tested in GSE41374, GSE152075, and GSE156063, and the locked blood module was tested in GSE171110 and the GSE38900 platform units GPL10558 and GPL6884. The table reports the numbers of positive and negative samples, represented module genes, AUROC, average precision, and Welch-test FDR. Internal anchor performance is omitted. AUROC and average precision are reported as portability summaries and not as estimates of clinical diagnostic performance.

Figure 4. External tissue-matched module-score portability. AUROCs are shown for the nasal module in GSE41374 (76 RSV, 10 controls), GSE152075 (430 SARS-CoV-2, 54 controls), and GSE156063 (93 SARS-CoV-2, 100 controls), and for the blood module in GSE171110 (44 SARS-CoV-2, 10 controls), GSE38900-GPL10558 (28 RSV, 8 controls), and GSE38900-GPL6884 (107 RSV, 31 controls). Scores are unweighted means of represented per-gene z values. The dashed line indicates AUROC = 0.5. Values are portability summaries, not clinical diagnostic estimates. Please click here to view a larger version of this figure.
Longitudinal validation tested whether scores fall during recovery
The GSE97741/GSE97742 companion datasets provided an independent natural-infection validation setting with acute and discharge samples from hospitalized children27. These samples were not used as healthy-control discovery data. They were used to assess whether anchor-derived module scores declined from acute illness to discharge.
For the prespecified combined RSV single-infection and rhinovirus group, 68 paired subjects were analysed in each tissue (Table 4; Figure 5). The blood module decreased from acute illness to discharge in blood (mean delta = 0.330; Cohen dz = 0.740; paired-test FDR = 1.98 × 10⁻7; AUROC = 0.765). The nasal module also decreased in nasopharyngeal samples (mean delta = 0.436; Cohen dz = 0.477; paired-test FDR = 3.06 × 10⁻4; AUROC = 0.679). Paired trajectories and their standard errors show that the group-level decline was not driven by a few unpaired extremes.
| Dataset | Sample source | Module | Tissue matched | n pairs | Mean acute-discharge delta | Cohen dz | AUROC | Paired-test FDR |
| GSE97741 | Blood | Blood | Yes | 68 | 0.330 | 0.740 | 0.765 | 1.98 × 10⁻⁷ |
| GSE97741 | Blood | Nasal | No | 68 | 0.479 | 0.721 | 0.755 | 3.19 × 10⁻⁷ |
| GSE97742 | Nasopharyngeal | Blood | No | 68 | 0.361 | 0.914 | 0.845 | 9.42 × 10⁻¹⁰ |
| GSE97742 | Nasopharyngeal | Nasal | Yes | 68 | 0.436 | 0.477 | 0.679 | 3.06 × 10⁻⁴ |
Table 4: Independent longitudinal acute-versus-discharge validation. The table reports the number of complete pairs, mean acute-minus-discharge score difference, Cohen dz, AUROC, and paired-test FDR for the fixed modules in GSE97741 and GSE97742. Positive delta indicates a higher module score during acute illness. Paired-test FDR values are Benjamini–Hochberg-adjusted two-sided paired t-test P values calculated across all 20 valid paired t tests in the full longitudinal output.

Figure 5. Paired module-score changes from acute illness to discharge. (A) Blood-module scores in whole blood (GSE97741). (B) Nasal-module scores in nasopharyngeal samples (GSE97742). Each panel contains 68 complete subject pairs: 38 RSV single infections and 30 rhinovirus infections. Thin lines connect subject-matched measurements. Orange points indicate group means, and orange error bars indicate the standard error of the mean. Two-sided paired t tests and Wilcoxon signed-rank tests were performed; Benjamini–Hochberg FDR correction was applied across the 20 valid longitudinal paired t tests. Please click here to view a larger version of this figure.
The cross-compartment tests revealed a useful nuance. The blood module applied to nasopharyngeal samples yielded an AUROC of 0.845, even though individual nasal and blood effect sizes were weakly concordant in the anchor. Inspection of paired trajectories showed a consistent score decline rather than a reversed label distribution. Thus, the result is not evidence that the same individual genes dominate both tissues. It indicates that a coordinated interferon/inflammatory program can be summarized by different but partially redundant gene sets. Compartment specificity is strongest at the gene-ranking level and is not absolute at the pathway or score level.
Held-out asymptomatic detections tested symptom-gradient behavior
Asymptomatic picornavirus detections in GSE117827 were excluded from module construction to avoid including them in the primary control group. This held-out group then provided a biological check. In both compartments, module scores increased across the ordered groups of virus-negative controls, asymptomatic picornavirus detection, and symptomatic infection (Figure 6A,B).

Figure 6. Module scores across the held-out symptom gradient. (A) Nasal module: 6 virus-negative controls, 5 asymptomatic picornavirus detections, and 15 symptomatic infections. (B) Blood module: 6 virus-negative controls, 5 asymptomatic picornavirus detections, and 13 symptomatic infections. Points represent individual samples. Centre lines indicate medians; boxes span the 25th to 75th percentiles; and whiskers extend to the most extreme values within 1.5 times the interquartile range. Labels report two-sided Mann–Whitney U post-hoc comparisons with Benjamini–Hochberg correction across three comparisons per compartment. Please click here to view a larger version of this figure.
The nasal module correlated with the ordinal symptom score (Spearman rho = 0.818, P = 3.28 × 10⁻7; Kruskal-Wallis P = 1.57 × 10⁻4). The blood module showed a similar gradient (rho = 0.861, P = 6.89 × 10⁻8; Kruskal-Wallis P = 1.92 × 10⁻4). After correction within each three-comparison family, symptomatic infection differed from controls and asymptomatic detections in both compartments. Asymptomatic detections did not differ from virus-negative controls (nasal FDR = 0.792; blood FDR = 0.082). The modules therefore tracked symptomatic host-response activity more clearly than viral detection alone.
Clinical benchmarks defined the boundary between host response and pathogen specificity
Clinical benchmarks defined the distinction between host-response activity and pathogen classification (Table 5; Figure 7). In GSE63990, the nasal module yielded AUROCs of 0.782 for viral versus bacterial illness and 0.791 for viral versus non-infectious illness. The blood module yielded AUROCs of 0.678 and 0.753, respectively. The Pandya 33-mRNA comparator performed better, with AUROCs of 0.867 and 0.852, respectively. This result is expected for a set designed for viral/non-viral discrimination and shows that the anchor modules should not be presented as replacement diagnostic classifiers.
| Dataset | Contrast | Anchor nasal | Anchor blood | Pandya 33 mRNA | Andres-Terre ISG | Hallmark interferon-alpha |
| GSE63990 | Viral vs bacterial | 0.782 | 0.678 | 0.867 | 0.833 | 0.831 |
| GSE63990 | Viral vs non-infectious | 0.791 | 0.753 | 0.852 | 0.851 | 0.848 |
| GSE40012 | Viral vs bacterial pneumonia | 0.755 | 0.789 | 0.893 | 0.867 | 0.872 |
| GSE40012 | Viral pneumonia vs SIRS | 0.897 | 0.907 | 0.985 | 0.965 | 0.956 |
| GSE40012 | Viral pneumonia vs healthy | 0.804 | 0.986 | 0.923 | 0.906 | 0.891 |
| GSE40012 | Bacterial pneumonia vs healthy | 0.476 | 0.917 | 0.465 | 0.391 | 0.378 |
| GSE53543 | Ex vivo rhinovirus-stimulated vs unstimulated PBMCs | 1 | 0.957 | 1 | 1 | 1 |
Table 5: AUROC benchmark for anchor modules and reference host-response sets. GSE63990 and GSE40012 are clinical whole-blood cohorts. GSE53543 is an ex vivo PBMC perturbation experiment involving 98 paired subjects and is reported separately from the natural clinical cohorts. Reference sets were scored as unweighted gene means rather than as their original weighted classifiers. AUROC values are reported as benchmarking metrics and not as estimates of clinical diagnostic performance.

Figure 7. AUROC benchmarking of anchor modules and reference host-response sets. Rows represent prespecified contrasts in GSE63990, GSE40012, and GSE53543; columns represent the two anchor modules and three unweighted reference sets. GSE53543 is labelled as ex vivo rhinovirus-stimulated versus unstimulated PBMCs and is presented separately from the natural clinical cohorts. The Hallmark comparator is labelled Hallmark interferon-alpha. Cell values indicate AUROCs used for method benchmarking and should not be interpreted as estimates of clinical diagnostic performance. Please click here to view a larger version of this figure.
GSE40012 sharpened this boundary. The Pandya comparator reached AUROCs of 0.893 for viral versus bacterial pneumonia and 0.985 for viral pneumonia versus SIRS. The blood module separated influenza A pneumonia from healthy controls (AUROC = 0.986) and from SIRS (AUROC = 0.907), but it also separated bacterial pneumonia from healthy controls (AUROC = 0.917). The blood module therefore measures broad systemic inflammatory and interferon-associated activity. It is not virus-specific, and a high score cannot assign pathogen class.
In the separate GSE53543 ex vivo PBMC benchmark, the nasal module and interferon comparators reached an AUROC of 1.000 for rhinovirus-stimulated versus media-only PBMCs; the blood module reached an AUROC of 0.957. All 98 subjects contributed paired conditions. This controlled result supports responsiveness of the scored programs to rhinovirus stimulation. It does not estimate clinical diagnostic performance.
Robustness checks tested module size, random alternatives, and selection stability
Anchor separation remained unchanged across 10, 25, 50, 100, and 200 top-ranked genes (Figure 8A). These internal AUROCs are not external validation, but they show that the qualitative result did not depend on choosing exactly 50 genes. In 500 random 50-gene sets, null AUROC medians were 0.489 for nasal and 0.705 for blood; the corresponding 99th percentiles were 0.722 and 0.872 (Figure 8B). The observed modules exceeded these null distributions. Bootstrap selection frequencies were distributed rather than concentrated in a single invariant list (Figure 8C). This finding directly reflects the small anchor sample. It supports a stable aggregate signal while cautioning against treating every selected gene as fixed.

Figure 8. Module-size, random-gene, and bootstrap robustness analyses. (A) Anchor AUROC across module sizes of 10, 25, 50, 100, and 200 genes; these values represent internal sensitivity checks. (B) AUROC distributions from 500 random protein-coding 50-gene sets sampled without replacement from genes represented in GSE117827 (seed = 20260622). Orange points indicate the observed module AUROCs. Horizontal lines within each violin indicate the 25th percentile, median, and 75th percentile. (C) Bootstrap reselection was restricted to the 1,000 positive-effect protein- coding genes ranked highest in the original anchor analysis for each compartment. Bars show the 20 genes with the highest selection frequency per compartment; selection frequency was calculated as the number of selections divided by 100. Please click here to view a larger version of this figure.
Marker-program and variance analyses clarified what the scores measure
Across GSE40012, GSE53543, and GSE63990, both modules correlated most consistently with the myeloid interferon program (Figure 9A). Nasal-module correlations were 0.812, 0.878, and 0.882; blood-module correlations were 0.517, 0.843, and 0.774. Variance partitioning was descriptive (Table 6; Figure 9B). For the blood module, condition and detailed group explained more variance (eta-squared = 0.282 and 0.250, respectively) than dataset (0.066), sample type (0.026), or source group (0.025). For the nasal module, detailed group and condition also explained more variance than dataset, sample type, or source group. Biological condition and detailed group therefore accounted for larger fractions of module-score variance than dataset or sample type, although non-zero dataset and composition effects remain a limitation of bulk public-data reuse.

Figure 9. Marker-program correlations and descriptive variance partitioning. (A) Spearman correlations between module scores and six marker-program scores in GSE40012, GSE53543, and GSE63990. Programs required at least three represented genes. P values were adjusted across all dataset-module-program correlations. Blank cells indicate unavailable or non-estimable combinations after gene and sample requirements. (B) One-way eta-squared (η2; between-group sum of squares/total sum of squares) for condition, detailed group, dataset, sample type, and source group. The analysis included 934 blood-module and 1,469 nasal-module scores and is descriptive, not causal. Please click here to view a larger version of this figure.
| Module | Factor | Eta-squared | n samples |
| Anchor blood | Condition | 0.282 | 934 |
| Anchor blood | Detailed group | 0.25 | 934 |
| Anchor blood | Dataset | 0.066 | 934 |
| Anchor blood | Sample type | 0.026 | 934 |
| Anchor blood | Source group | 0.025 | 934 |
| Anchor nasal | Detailed group | 0.17 | 1469 |
| Anchor nasal | Condition | 0.13 | 1469 |
| Anchor nasal | Dataset | 0.021 | 1469 |
| Anchor nasal | Sample type | 0.006 | 1469 |
| Anchor nasal | Source group | 0.002 | 1469 |
Table 6: Descriptive variance partitioning of module scores. One row is provided for each module-factor pair, with the corresponding eta-squared and sample size. Eta-squared was calculated as the between-group sum of squares divided by the total sum of squares after excluding rows missing the relevant module score or factor. Factors were evaluated individually; therefore, the analysis is descriptive, does not adjust for mutually correlated factors, and should not be interpreted causally.
Data Availability:
All transcriptomic datasets analyzed in this study are publicly available from the Gene Expression Omnibus under accession numbers GSE117827, GSE41374, GSE152075, GSE156063, GSE171110, GSE38900, GSE97741, GSE97742, GSE63990, GSE40012, and GSE53543. The analysis-derived data underlying the main figures and tables and the analysis code are available from the corresponding author upon reasonable request. No restricted or newly generated participant-level data were used in this study.