The protocol was validated using 315 Salmonella isolates from Southwest China, with 213 isolates (2/3) serving as the training set to compare five WGS-based serotype prediction methods (MLST, SISTR, SeqSero, SeqSero2, and SeqSero2S) against conventional serotyping results, while the remaining 102 isolates (1/3) were used as a validation set to assess MLST, SISTR, and SeqSero2S performance.
Training set analysis (n = 213)
Reproducibility assessment: Conventional serotyping demonstrated limited concordance with any of the five WGS-based prediction methods, achieving only 60.6% agreement (129/213). In contrast, pairwise agreement among any two of the five WGS approaches reached 99.1% (211/213), and complete five-method consensus was observed in 82.2% of isolates (175/213). This substantial disparity (82.2% vs. 60.6%, P<0.001) underscores the superior reproducibility of genomic approaches compared to conventional serology (Figure 1).
Concordance ranking: Comparison among WGS prediction methods showed the following descending order of concordance rates: SISTR (100%, 213/213), MLST (98.1%, 209/213), SeqSero2S (90.1%, 192/213), and SeqSero2/SeqSero (both 83.1%, 177/213). While no significant difference existed between MLST and SISTR (P=0.123), both MLST and SISTR showed significantly higher concordance than SeqSero/SeqSero2/SeqSero2S (P<0.001). Furthermore, SeqSero2S demonstrated statistically superior performance relative to both SeqSero and SeqSero2 (P=0.033), establishing a clear performance gradient: SISTR> MLST> SeqSero2S>SeqSero2≈SeqSero (Figure 1, Supplementary File 1).
Validation set confirmation (n = 102)
Reproducibility assessment: Concordance between conventional serotyping and any WGS methods was 66.7% (68/102), while pairwise WGS agreement reached 100% (102/102), with five-method consensus in 88.2% (90/102), reaffirming the superior reproducibility of WGS-based methods (88.2% vs. 67.6%, P=0.004) (Figure 2).
Concordance ranking: WGS methods showed similar concordance: SISTR (100.0%, 102/102), MLST (97.1%, 99/102), SeqSero2S (93.1%, 95/102), and SeqSero2/SeqSero (both 91.2%,93/102). While MLST and SISTR showed no significant difference (P=0.246), only SISTR showed significantly higher concordance than SeqSero/SeqSero2/SeqSero2S (P<0.05). The same performance hierarchy was observed: SISTR> MLST> SeqSero2S>SeqSero2≈SeqSero (Figure 2, Supplementary File 1).
Total set evaluation (n = 315)
In the comprehensive evaluation of all 315 samples, SISTR demonstrated complete coverage (100%, 315/315). MLST and SeqSero2S also showed high concordance rates of 97.8% (308/315) and 91.1% (287/315), respectively. The superior coverage of genomic methods contrasted sharply with traditional serotyping (84.1% (265/315) vs. 62.9% (198/315), P<0.001), highlighting its limitations as a primary surveillance tool (Figure 3A).
To systematically evaluate method performance, we assessed pairwise concordance across all 315 isolates. As shown in the inter-method concordance heatmap (Figure 3B), genomic approaches exhibited exceptional mutual agreement, with consistency scores exceeding 0.89 among SISTR, MLST, and SeqSero2S. In stark contrast, traditional serotyping showed substantially lower concordance (0.49-0.57) with all genomic platforms, revealing a fundamental methodological divergence.
The consensus network analysis (Figure 3C) further confirmed these relationships, showing genomic methods forming a tightly interconnected core with SISTR and SeqSero2S as central hubs. The robust connections among genomic tools underscore their reliability, while the weak links connecting traditional serotyping visually emphasize its inconsistent performance.
For detailed serotype-level analysis, we used SISTR predictions as the benchmark due to its complete coverage and examined concordance with other methods in Table 1. Among the 15 major Salmonella serotypes (n≥5), MLST demonstrated perfect concordance (100%) with SISTR for 13/15 serotypes, while SeqSero2S achieved 100% concordance for 11/15 serotypes. Remarkably, conventional serotyping failed to achieve perfect concordance with SISTR for any of the 15 major serotypes (0/15), with concordance rates ranging from 0% to 95.0% (Table 1, Supplementary File 1).
Interpretation guidelines
For reliable prediction: Consistent with the high inter-method agreement observed among genomic tools, a consensus among SISTR, MLST, and Seqsero2S (concordance ≥90% in this study) provides a highly reliable serotype prediction.
For tool selection: The performance hierarchy (SISTR ≈ MLST > SeqSero2S) provides a data-driven basis for selecting tools based on required accuracy and resources, as detailed in the recommended workflow (section 6).
Suboptimal case
Genomic vs. traditional serotyping discrepancies: The low concordance (0.49-0.57) between genomic platforms and traditional serotyping reflects a fundamental methodological divergence. In cases of discrepancy, the consensus from genomic methods (especially SISTR and MLST) should be prioritized, as they demonstrate superior reproducibility and coverage.

Figure 1: UpSet plot comparing traditional serotyping with five WGS-based prediction methods (MLST, SeqSero, SeqSero2, SeqSero2S, SISTR) in the test set (n = 213). This UpSet plot visualizes the agreement and discrepancies between conventional serotyping and the five whole-genome sequencing (WGS)-based prediction methods (MLST, SeqSero, SeqSero2, SeqSero2S, SISTR). The horizontal bar chart (left) shows the size of each set (the number of isolates in agreement assigned to each method). The vertical bar chart (top) shows the size of the intersections between methods (the number of isolates where predictions agreed). The dot and line matrix (bottom) indicates which methods are included in each intersection. In the test set (n = 213), significantly higher inter-method agreement among the five WGS-based approaches (top3: SISTR, MLST, SeqSero2S) compared to their concordance with traditional serotyping. Please click here to view a larger version of this figure.

Figure 2: UpSet plot validating serotyping against five WGS-based methods (MLST, SeqSero, SeqSero2, SeqSero2S, SISTR) in the verification set (n = 102). This UpSet plot validates the findings from the training set. The results reaffirm the superior reproducibility of WGS-based methods, as indicated by the large intersection among SISTR, MLST, and SeqSero2S, compared to their smaller shared intersections with traditional serotyping. Please click here to view a larger version of this figure.

Figure 3: Comparison of serotype prediction concordance between traditional serotyping and five in silico genomic methods (n = 315). (A) UpSet plot of method agreement patterns. SISTR achieves complete coverage (315/315, 100%), followed by MLST (308/315, 97.8%), SeqSero2S (287/315, 91.1%), SeqSero2 (270/315, 85.7%), SeqSero (270/315, 85.7%), and traditional serotyping (198/315, 62.9%). (B) Heatmap of inter-method concordance. Genomic methods (MLST, SISTR, SeqSero2S) demonstrate high mutual agreement (>0.89), whereas traditional serotyping shows substantially lower concordance with these genomic methods (0.49-0.57). The strong concordance among genomic platforms supports their reliability for Salmonella surveillance. (C) Consensus network of serotyping approaches. This network visualization depicts concordance relationships among the six serotyping approaches, where edge thickness and color intensity represent the level of agreement. The circular layout reveals two distinct clusters: genomic methods form a tightly interconnected core with strong consistency (>0.84), while traditional serotyping remains peripherally connected with weaker links (0.49-0.57). SeqSero2S shows robust connections to all genomic methods (>0.90). Please click here to view a larger version of this figure.
Table 1: Concordance between SISTR-predicted serotypes and conventional serotyping/MLST/ SeqSero2S results for Salmonella strains (n = 315). Please click here to download this Table.
Supplementary File 1: Clean data. Please click here to download this File.