Method Article

Source-Audited Mining and Visualization of Real-World Traditional Chinese Medicine Case Records Using Hegan Jianpi Formula

10 views

⸱

DOI:

10.3791/73716

⸱

September 29th, 2026

* These authors contributed equally

In This Article

Summary

This protocol integrates de-identification, terminology normalization, source verification, and reproducible visualization to generate auditable, aggregate outputs from real-world traditional Chinese medicine outpatient records.

Abstract

Real-world traditional Chinese medicine (TCM) outpatient records contain longitudinal information on diagnoses, syndrome elements, symptom narratives, and individualized prescriptions, but their semi-structured format complicates reproducible analysis. This article presents a source-audited protocol for converting de-identified case records into traceable aggregate tables and publication-ready visualizations. The workflow defines eligibility criteria and analysis units, preserves mappings between original and normalized terminology, verifies all reported numerators and denominators, and regenerates figures from versioned source tables. Applied to records from January 2015 through June 2024, the protocol identified 555 eligible prescription visits from 396 patients. The final dataset contained 324 normalized herb labels and 9,339 herb-use occurrences. The workflow generated parallel Western medicine and TCM disease spectra, an explicit syndrome-element spectrum and co-occurrence matrix, herb-frequency and materia medica profiles, a 37-term history-derived symptom spectrum, 15 directed association rules, hierarchical clustering, and diagnosis-stratified herb-use heatmaps. At least one prespecified history keyword was detected in 474 visits (85.41%). When permitted by institutional policy, a human-supervised artificial intelligence (AI) option may assist with terminology checks, cross-file consistency review, and caption-figure alignment. Any use of AI must be documented in a separate audit record and must not replace source-data review. By integrating governance, normalization, computation, clinical interpretation, and visual output into a single reproducible workflow, the method transforms dispersed clinical documentation into a reviewable research resource. The protocol can be adapted to other expert-practice datasets, multicenter terminology projects, and prospective clinical data platforms.

Introduction

Real-world traditional Chinese medicine (TCM) outpatient records preserve longitudinal clinical reasoning through narrative symptoms, parallel diagnostic systems, syndrome descriptions, and individualized prescriptions. Early knowledge-discovery work recognized case records as complex experiential data requiring dedicated informatics methods before their patterns can be studied reliably1. Subsequent research has shown that synonymous symptom expressions, fine-grained entities, and locally documented diagnostic terms must be normalized without erasing the original wording2,3,4,5,6. Electronic health record (EHR) data can also support computational prescription research and the use of reusable analytical tools when clinical context and data lineage remain visible7,8. Authoritative materia-medica references are therefore needed to standardize herb names, processing forms, properties, flavors, and meridian assignments9,10. At the reporting level, the RECORD and STROBE statements call for transparent descriptions of data sources, eligibility rules, variables, and analytical decisions, while the FAIR principles emphasize traceable and reusable research objects11,12,13. These considerations are especially important when records contain repeated visits, missing fields, overlapping terminology, and free text. Fitness-for-use assessment must address completeness, conformance, plausibility, and source agreement14,15. Free-text de-identification also requires explicit procedures and human review because automated methods can preserve useful clinical content while leaving residual privacy risks16,17,18.

Once governance and terminology are established, frequency profiling, association-rule mining, and hierarchical clustering can provide complementary views of prescription structure19,20,21. Network and science-mapping methods further demonstrate how coordinated visual representations can reveal relationships that a single ranked list cannot capture22,23,24,25. Recent studies of rheumatoid arthritis with depression, Janus kinase inhibitors in ulcerative colitis, and rheumatoid arthritis with fibromyalgia illustrate the value of version-locked retrieval criteria, co-occurrence networks, clustering, temporal interpretation, and explicit statements of analytical limitations26,27,28. The present study adapts these transferable design principles to clinical record mining rather than to literature bibliometrics. In the source records, Hegan Jianpi Formula refers to an experience-based formula associated with Professor Nai-li Yao's practice29,30. Related publications describe his broader liver-spleen harmonization and spleen-stomach clinical approach29,30. In this method article, the formula name identifies the source-record context only and does not establish a fixed composition, dose, treatment indication, or efficacy claim. The computational workflow uses deterministic rules and general-purpose statistical methods. Reproducible computational practice additionally requires preservation of inputs, documentation of parameters, scripted transformations, and reviewable outputs31,32. A human-supervised artificial intelligence (AI) option can assist with terminology checks, cross-file consistency reviews, and visual quality control when its use is documented. Clinical expertise, privacy safeguards, and accountable human decision-making must remain primary (Supplementary Table 1)33,34. The overall goal of this method is to convert de-identified TCM case records into a source-audited chain of disease-, syndrome-, symptom-, herb-use-, association-, clustering-, and diagnosis-stratified outputs. The worked example uses 555 eligible prescription visits from 396 patients to demonstrate how governance, normalization, computation, clinical interpretation, and visual release can be integrated without disclosing patient-level records, exact formulation details, dose ratios, or prescriptive instructions.

Protocol

This retrospective analysis of de-identified clinical records was reviewed and approved by the Medical Ethics Committee of Guang'anmen Hospital, China Academy of Chinese Medical Sciences (approval No. 2026-108-KY; 13 July 2026). The requirement for informed consent was waived by the Committee.

1. Establish ethics, governance, and data security

  1. Assign a data custodian, source-data analyst, clinical terminology reviewer, quantitative reviewer, and investigator authorized to approve aggregate release. Record each role in the project log.
  2. Prepare a data-management plan specifying the permitted source system, study period, fields, storage location, access permissions, backup procedure, retention period, and linkage-key restrictions.
  3. Export only the minimum variables required for the prespecified analysis. Include a local patient identifier, visit date, age or date of birth when permitted, sex, Western diagnosis, TCM disease name, de-identified history text, tongue text, pulse text, syndrome text, and prescription-herb fields.
  4. Remove names, national identification numbers, telephone numbers, detailed addresses, insurance identifiers, contact information, and other direct identifiers within the controlled institutional environment. Replace each medical record number with a project-specific patient identifier.
  5. Scan free-text fields for residual identifiers and mask every detected item. Store any linkage key separately from the analytical dataset and exclude it from working, collaboration, and submission folders.
  6. Preserve the initial de-identified export as a read-only source snapshot. Record its filename, creation date, row count, column count, file size, and checksum in a source manifest.
  7. Create separate directories for source files, working files, dictionaries, aggregate tables, figures, scripts, and audit records.
  8. Restrict released or shared materials to aggregate tables, approved de-identified dictionaries, scripts, and publication figures. Do not upload identifiable or potentially re-identifiable records to public generative-AI systems, unapproved cloud services, or public repositories.
    NOTE: When institutional policy permits human-supervised AI assistance, provide only de-identified text excerpts or aggregate tables. Record the purpose, reviewed output, accepted correction, reviewer, and date in the audit log.

2. Define the record-identification rule and analysis units

  1. Define the institution, outpatient setting, responsible clinical team, electronic medical record system, and source period before screening. For the worked example, use records dated January 2015 through June 2024.
  2. Store the cohort-identification rule in a restricted, version-controlled file before examining analytical results. Record accepted terminology variants and all inclusion and exclusion conditions without disclosing confidential formulation details in the manuscript.
  3. Apply the version-locked record-identification rule at the prescription-visit level after herb-name normalization. Finalize the source files, dictionaries, cohort rule, and analysis parameters as a single release set.
  4. Create a new version and regenerate all dependent outputs whenever a version-locked component changes. Do not modify the cohort-identification rule after examining disease distributions, herb frequencies, association rules, or clusters.
  5. Include an outpatient visit when it falls within the version-locked study period, maps to a valid project patient identifier, contains an interpretable prescription record, and satisfies the restricted record-identification rule.
  6. Exclude exact duplicate exports, records outside the study period, records without an interpretable prescription, and rows that fail the version-locked identification rule. Assign one primary exclusion reason to each excluded row.
  7. Distinguish repeated exports from genuine follow-up visits. Remove only exact system or export duplicates and retain genuine repeated visits.
  8. Assign stable patient and prescription-visit identifiers. Create a cohort manifest containing the source version, rule version, screening status, inclusion or exclusion reason, patient identifier, visit identifier, and reviewer status.
  9. Use one record per unique patient for fixed demographic summaries, including age and sex. Use the prescription visit as the analysis unit for disease, syndrome, history-keyword, tongue, pulse, herb, association-rule, and clustering analyses.
  10. Separate cohort eligibility from variable availability. Retain an otherwise eligible visit when a variable-specific field is missing and report the evaluable denominator for each analysis.
  11. Freeze the cohort manifest before descriptive analysis. Record the number of eligible prescription visits and unique patients, and regenerate all dependent outputs if an eligibility decision changes.

3. Normalize terminology and preserve the audit trail

  1. Create separate version-controlled dictionaries for herbs, Western diagnoses, TCM disease names, syndrome elements, history keywords, tongue terms, pulse terms, and materia-medica attributes.
  2. Include the original term, normalized term, term category, source field, rule, reference source, dictionary version, reviewer, review date, and comments in each dictionary. Retain every original term after normalization.
  3. Normalize Chinese medicinal material (CMM) names using authoritative nomenclature references9,10. Correct typographical variants and system abbreviations while preserving clinically meaningful processing forms.
  4. Retain processing qualifiers, including stir-fried, vinegar-processed, honey-fried, ginger-processed, wine-processed, charred, and raw, as separate labels.
  5. Keep orthographically or clinically distinct herbs separate. Do not merge Bai Shao with Bai Zhu or infer an herb from an ambiguous abbreviation without source review.
  6. Preserve the original dose and unit in separate fields when available. Remove dose text only when constructing visit-level herb-presence variables.
  7. Do not interpret occurrence frequency as cumulative dose or treatment intensity.
  8. Normalize Western and TCM disease names independently. Preserve the diagnostic system and principal-diagnosis status, and do not translate a TCM disease label into a Western diagnosis unless both are explicitly recorded.
  9. Define a transliterated TCM disease label at first use when needed for reader orientation. Retain the original source label.
  10. Split multivalued explicit syndrome text using version-locked delimiters and map source expressions to normalized syndrome elements. Deduplicate each normalized element within a visit.
  11. Keep explicit syndrome text separate from inherited or score-derived indicator columns.
  12. Retain the original tongue, pulse, structured symptom, and de-identified history fields as separate data products. Do not treat a history-keyword match as equivalent to a clinician-entered structured symptom.
  13. Ask one reviewer to propose each ambiguous or many-to-one mapping, and have a second clinical reviewer verify it. Exclude uncertain mappings from released analyses until source review resolves them.
  14. Check for blank normalized labels, one-to-many mappings, many-to-one mappings, duplicate dictionary entries, unreviewed terms, and accidental loss of processing labels. Freeze dictionary versions before generating aggregate tables.

4. Verify source fields and lock reportable outputs

  1. Create an evidence ledger with one row for each manuscript claim, table, and figure panel. Record the analysis unit, source field, source version, dictionary version, script version, numerator, denominator, output file, reviewer, and approved wording.
  2. Recalculate eligible-visit and unique-patient counts, demographic summaries, disease counts, explicit syndrome counts, history-keyword counts, merged-herb frequencies, and herb-use totals directly from the version-locked files.
  3. Compare each recalculated value with the corresponding manuscript, spreadsheet, or figure value. Correct the manuscript and regenerate dependent outputs whenever a discrepancy is confirmed.
  4. Resolve each discrepancy at the source-field level. Record the reason, correction, reviewer, and date in the evidence ledger before releasing the affected result.
  5. Verify that each aggregate table contains the normalized label, numerator, denominator, percentage definition, analysis unit, and source identifier required to reproduce the corresponding statement.
  6. Verify the denominator used for each percentage. Use unique patients for fixed demographic characteristics and prescription visits for dynamic clinical and prescription variables.
  7. Compare figure labels, values, ordering, and panel definitions with their version-locked aggregate tables. Regenerate each figure after a table correction rather than manually editing plotted values.
  8. Ask a clinical reviewer to verify terminology mappings. Ask a quantitative reviewer to verify counts, rates, rule metrics, and heatmap denominators.
  9. Lock the evidence ledger, aggregate tables, and figure sources under one version identifier before manuscript assembly.

5. Generate descriptive spectra and co-occurrence outputs

  1. Generate patient-level age and sex summaries from one row per unique patient. Retain and explicitly report unknown categories.
  2. Generate Western and TCM disease spectra separately at the prescription-visit level. Preserve the original diagnostic system and calculate leading normalized labels using a common visit denominator.
  3. Generate the syndrome-element spectrum from the explicit normalized syndrome-text field. Split terms using version-locked delimiters, deduplicate each term within a visit, and calculate visit-level frequencies.
  4. Build the syndrome co-occurrence matrix from the same within-visit token sets. Encode node size by visit frequency and edge width or heatmap intensity by pairwise co-occurrence count.
  5. Prespecify an exact history-keyword dictionary when symptom information is primarily narrative. Scan only the de-identified history field and count each concept no more than once within a visit.
  6. Calculate the proportion of eligible visits containing at least one history keyword and rank all keyword counts. Export the full table and the plotted subset.
  7. Generate merged-herb frequencies from normalized prescription data. Count each normalized herb label no more than once per visit, and retain clinically meaningful processing forms as separate labels.
  8. Generate Four-Qi, five-flavor, and meridian-tropism profiles from the version-locked materia-medica dictionary. Treat Four-Qi as a single-valued category for each property-coded herb occurrence.
  9. Treat five-flavor and meridian-tropism as multi-label attributes. State the property-coded denominator and preserve the separate analysis unit.
  10. Export one machine-readable aggregate table for each descriptive domain before generating the final figures.

6. Perform association-rule and hierarchical-clustering analyses

  1. Construct a binary prescription-visit-by-herb matrix from the version-locked normalized prescription table. Use one row per eligible visit and one column per normalized herb label.
  2. Calculate support, confidence, and lift for directed single-herb association rules19. Retain the direction of each rule because confidence depends on the antecedent.
  3. Apply the prespecified thresholds of support ≥ 0.30, confidence ≥ 0.70, and lift > 1.0. Export each retained rule with its co-occurrence count and all three metrics.
  4. Draw the complete retained rule set as a directed bipartite rule network and an ordered rule-strength profile. Encode network edge width by support and edge color by lift.
  5. Encode point size by support and annotate confidence in the ordered profile.
  6. Verify that both rule displays contain the same retained rule set and preserve rule direction. Reconcile each mismatch against the exported rule table before release.
  7. Select the 20 highest-frequency herb-presence variables for hierarchical clustering. Calculate pairwise Jaccard distances and apply average linkage.
  8. Label the dendrogram with normalized herb names. Interpret branch proximity only as similarity in visit-level occurrence patterns.
  9. Export the binary matrix specification, rule table, network edge list, clustering input order, distance matrix, and linkage matrix before figure generation.
  10. Ask a quantitative reviewer to reproduce one rule metric manually. Compare every network edge with the rule table.
  11. Rerun the complete rule and clustering workflow whenever the cohort, terminology dictionary, or herb-presence matrix changes.
  12. For the worked example, perform a deterministic one-visit-per-patient sensitivity analysis by retaining the earliest source-order record for each patient. Compare the resulting 12 rules with the 15 rules from the visit-level analysis.
  13. Report the change as a descriptive robustness result rather than patient-level validation. Use the deterministic recomputation provided in Supplementary File 1.

7. Generate diagnosis-stratified patterns and the clinical interpretation framework

  1. Select clinically relevant principal Western-diagnosis strata before inspecting diagnosis-specific herb distributions.
  2. Count each normalized herb label no more than once per visit within each diagnosis stratum. Calculate prevalence as the number of visits containing the label divided by the total number of eligible visits in that stratum.
  3. Use the same displayed herb set and a fixed 0%–100% color scale across diagnosis strata. Display cell values to preserve exact interpretation.
  4. Report the numerator and denominator for each displayed cell in an independent aggregate table.
  5. Build a separate clinical interpretation panel after completing the descriptive heatmap. Keep computed results and expert interpretation visually distinct.
  6. Populate the cue layer only with terms traceable to the version-locked history-keyword table or explicit normalized syndrome text.
  7. Ask two clinical reviewers to agree on concise contextual interpretations of the selected cues. Record the reviewers and final wording in the evidence ledger.
  8. Link each contextual interpretation to a research use, such as variable selection, prospective validation, or hypothesis generation.
  9. Use dashed, non-directional connectors for the interpretation layer. Omit dose, fixed-composition, and treatment-recommendation content.
  10. Review the heatmap and interpretation panel together. Ensure that computed prevalence, expert context, and future research use remain clearly separated.

8. Assemble and verify the reproducibility package

  1. Export the source manifest, cohort manifest, dictionary versions, evidence ledger, analysis scripts, aggregate tables, figure sources, and final figures into separate directories.
  2. Name each file using a stable figure or table number and version date. Preserve one authoritative current version and archive superseded outputs separately.
  3. Generate each figure as one multipanel image file. Export a high-resolution PDF and a 300 dpi PNG for review.
  4. Place figure legends in the manuscript and keep them outside the image files. Describe each panel, analysis unit, denominator, and visual encoding.
  5. Prepare one independent XLSX file for each table.
  6. Ask an independent reviewer to compare manuscript numbers with aggregate tables and percentages with their numerators and denominators.
  7. Compare rule edges with the rule table and heatmap cells with the source matrix.
  8. Run automated checks for filename consistency, required sections, figure dimensions, table count, reference count, template residue, and restricted wording.
  9. Ask the clinical author team to review terminology, interpretation, intellectual-property boundaries, author details, funding order, ethics wording, and the final figure set.
  10. Remove patient-level files, linkage keys, unrestricted free text, internal correspondence, and temporary review artifacts from the submission package.
  11. Freeze the released package and record its version, date, checksum, manuscript version, and author-approval status. Create a new version and a change log for each subsequent correction.

Results

Successful application of the protocol produced a traceable chain connecting the de-identified source snapshot, cohort manifest, terminology dictionaries, aggregate tables, scripts, and final figures. The version-locked screening procedure identified 555 eligible prescription visits from 396 unique patients. Figure 1 summarizes the three coordinated stages of the method: governance, terminology normalization, and analysis with clinical review. The corresponding transparency and audit documentation, including the aggregate-only boundary for AI-assisted review, is provided in Supplementary Table 1.

The final cohort comprised 555 eligible prescription visits contributed by 396 unique patients (Figure 2 and Table 1). Genuine repeated visits were retained for dynamic clinical and prescription variables, whereas age and sex were summarized after patient-level deduplication. Patient age ranged from 13 to 94 years, with a mean of 47.11 years and a standard deviation of 14.88 years. The cohort included 228 female patients (57.58%), 167 male patients (42.17%), and 1 patient with unknown sex (0.25%). This separation of analysis units preserved longitudinal encounters without inflating the demographic denominator.

Western and TCM disease spectra are summarized in Figure 3. Western diagnoses were concentrated in gastrointestinal, hepatobiliary, and pancreatic disorders. Hierarchical grouping assigned 271 visits to gastrointestinal disorders, 222 to hepatobiliary or pancreatic disorders, and 62 to other systems (Figure 3A). The leading individual diagnoses were chronic gastritis in 133 visits (23.96%), fatty liver in 77 visits (13.87%), and liver cirrhosis in 46 visits (8.29%), followed by chronic atrophic gastritis, abdominal pain, chronic hepatitis B, dyspepsia, reflux esophagitis, chronic superficial gastritis, and chronic pancreatitis (Figure 3C and Table 2).

The corresponding TCM disease-name spectrum organized the same visits according to a parallel diagnostic system. Spleen-stomach or intestinal disease names accounted for 280 visits, liver-gallbladder disease names for 223, qi-blood-fluid disorders for 26, mind-chest-head disorders for 12, channel or limb disorders for 5, gynecological disorders for 4, and other categories for 5 visits (Figure 3B). Wei pi, a retained TCM disease label for epigastric fullness or discomfort, was recorded in 163 visits (29.37%). Ji disease (an accumulation-type source TCM label), Gan zhuo (literally, liver fixity; a retained classical TCM disease label), Bi disease (impediment disease), and Xiao ke (wasting-thirst) were retained as source TCM labels and were not converted into modern biomedical diagnoses. Gan pi, a TCM disease label used in the contemporary Chinese consensus for fatty liver disorders, was recorded in 77 visits (13.87%)35. The Western and TCM source fields remained independent and were not substituted for one another (Figure 3D and Table 3). Xie xie (loose or watery stools) and Fu xie (diarrhea) were separate original TCM disease labels, recorded 6 and 2 times, respectively, and neither was equated with a modern biomedical diagnosis. Displaying the two spectra side by side preserved the logic of each diagnostic system while demonstrating the clinical breadth of the source records.

The explicit normalized syndrome-text field contained 555 eligible visit rows. After within-visit deduplication, the leading elements were spleen in 362 visits (65.23%), liver in 304 (54.77%), qi deficiency in 295 (53.15%), stomach in 171 (30.81%), dampness in 151 (27.21%), and qi stagnation in 109 (19.64%; Figure 4A and Table 4). Heat, water retention, blood, blood stasis, channels, kidney, heart, and yin deficiency comprised the remainder of the displayed spectrum. Percentages were nonexclusive because multiple syndrome elements could occur during a single visit.

The network and count matrix were generated from the same visit-level token sets (Figure 4B,C). Frequent co-recording centered on spleen, liver, and qi deficiency, with additional connections to stomach, dampness, qi stagnation, heat, and water retention. Using one version-locked matrix for the frequency plot, network, and heatmap ensured that all three panels represented the same underlying evidence at different levels of detail. The recovered audit identified 0 exact syndrome-text/indicator matches among 555 rows and a mean text-indicator Jaccard similarity of 0.1806. The indicator columns were therefore excluded from substantive reporting. The retained materials did not contain reconstructable pre-cleaning distribution vectors; consequently, cross-version distributional stability was not estimated. The one-visit-per-patient sensitivity analysis instead addressed the effect of repeated visits. Source audit and data quality outcomes are summarized in Supplementary Table 2.

After terminology normalization and retention of clinically meaningful processing forms, the 555 visits contained 324 distinct merged herb labels and 9,339 herb-use occurrences. Bai Shao was recorded in 531 visits (95.68%), followed by Fu Ling in 520 (93.69%), Dang Gui in 510 (91.89%), Chi Shao in 467 (84.14%), and stir-fried Bai Zhu in 361 (65.05%; Figure 5A and Table 5). The rank-frequency curve of 36 leading labels showed a steep, high-frequency segment followed by a progressively broader tail (Figure 5B), providing a compact representation of the shared-background and individualized-prescription components. The source-to-normalized mappings, processing qualifiers, standard CMM names, source species, and evidence boundaries underlying herb nomenclature are documented in Supplementary Table 3.

The auxiliary materia-medica dictionary contained 9,115 property-coded herb-use occurrences. The leading Four-Qi categories were warm (2,588, 28.39%), slightly cold (2,339, 25.66%), neutral (2,330, 25.56%), and cold (1,221, 13.40%; Figure 5C). After normalization of compound source flavor terminology, the leading five-flavor categories were bitter (4,488, 49.24%), sweet (4,377, 48.02%), and pungent (2,893, 31.74%) (Figure 5D). Meridian-tropism coding was led by spleen (5,541, 60.79%), liver (4,801, 52.67%), lung (3,358, 36.84%), stomach (2,988, 32.78%), and heart (2,702, 29.64%; Figure 5E). Figure 5F records the distinct visit-level and property-coded analysis units and their reproducibility link. Together, these panels demonstrate how prescription frequency and standardized materia-medica descriptors can be summarized without collapsing processing-specific terminology.

The prespecified history dictionary contained 37 exact symptom concepts. At least one concept was detected in 474 of 555 eligible visits (85.41%). The leading terms were dry throat in 178 visits (32.07%), fatigue in 151 (27.21%), poor sleep in 139 (25.05%), abdominal distension in 115 (20.72%), acid reflux in 100 (18.02%), bitter taste in 99 (17.84%), heartburn in 90 (16.22%), hiccup in 87 (15.68%), belching in 87 (15.68%), and loose stool in 77 (13.87%; Figure 6A,B and Table 6).

The complete rank-frequency curve displayed the full 37-term distribution, whereas the cumulative line summarized the concentration of all keyword detections across ranks (Figure 6C). Figure 6D records the reproducible sequence used to generate the spectrum: lock the dictionary, scan de-identified history text, deduplicate concepts within each visit, and export aggregate counts and rates.

Association-rule mining used the version-locked 555-row prescription-visit-by-herb presence matrix. Fifteen directed single-herb rules met support ≥ 0.30, confidence ≥ 0.70, and lift > 1.0 (Figure 7A,B and Table 7). Retained support values ranged from 0.3009 to 0.7892, confidence values from 0.7302 to 0.9748, and lift values from 1.0010 to 1.1226.

The most frequently co-recorded directed pair was Chi Shao > Fu Ling, based on 438 visits, with support of 0.7892, confidence of 0.9379, and lift of 1.0010. The reverse rule had the same support and a confidence of 0.8423, demonstrating why rule direction must be retained. Threshold sensitivity retained 18, 11, 21, 14, 7, and 4 rules under support ≥ 0.25 or ≥ 0.35, confidence ≥ 0.65 or ≥ 0.75, or lift > 1.02 or > 1.05, respectively. In the one-visit-per-patient analysis, the same pair occurred in 316 visits, with support of 0.7980, confidence of 0.9489, and lift of 0.9994; therefore, it was not retained under the lift > 1.0 criterion. Threshold perturbations and repeat-visit sensitivity results are provided in Supplementary Table 4, and the deterministic recomputation is provided in Supplementary File 1. Accordingly, high visit-level support should not be interpreted as patient-level validation.

The largest lift values were observed for Gan Cao > stir-fried Bai Zhu (1.1226) and Gan Cao > Chi Shao (1.1200). Figure 7A integrates all 15 retained rules in a directed bipartite network, with edge width encoding support and edge color encoding lift. Figure 7B ranks the same rules by lift, scales each point by support, and annotates confidence. The two panels, therefore, distinguish network topology from quantitative rule strength rather than duplicating the same visual representation.

Hierarchical clustering of the 20 highest-frequency herb variables used Jaccard distance and average linkage (Figure 7C). The dendrogram provides a global representation of similarity in visit-level presence patterns that complements the local directional information provided by the association rules. Together, the panels demonstrate how complementary analytical views can be aligned to a single version-locked binary matrix.

Diagnosis-stratified prevalence was calculated for fatty liver (77 visits), liver cirrhosis (46 visits), and chronic gastritis (133 visits; Figure 8A and Table 8). Bai Shao, Fu Ling, and Dang Gui showed high prevalence across all three strata. Liver-cirrhosis visits showed prominent documentation of Dan Shen (95.65%), vinegar-processed E Zhu (86.96%), vinegar-processed Bie Jia (84.78%), and Sheng Huang Qi (60.87%). Fatty-liver visits showed a higher prevalence of Yin Chen (41.56%) and Ze Xie (31.17%) than the other strata displayed, whereas chronic-gastritis visits showed recurrent Huang Lian (44.36%), Mu Xiang (34.59%), and Chai Hu (36.09%). The heatmap, therefore, provides a compact descriptive comparison of aggregate prescription patterns across major clinical contexts.

Figure 8B separates computation from interpretation. Verified history cues are linked to concise expert contexts and then to research uses, including contextualizing syndrome elements, comparing diagnosis-stratified patterns, selecting variables for prospective validation, and generating follow-up questions. This final panel demonstrates how clinical expertise can enrich the analytical output while remaining visibly distinct from computed prevalence.

Figure 8C summarizes the reproducibility and release boundary by distinguishing source-verified or reanalyzed components from elements requiring caution or further confirmation. This final audit layer prevents unresolved or nonreconstructable analytical products from being presented as validated outputs.

figure-results-1
Figure 1: Source-audited case-record mining workflow. Governance, terminology normalization, aggregate analysis, clinical review, and release packaging are shown as sequential stages of the workflow. Blue, green, and orange bands distinguish workflow domains, and dashed vertical lines indicate audit gates. Arrows indicate workflow sequence and do not imply biological causality. The release rule specifies that only aggregate outputs are released, patient identifiers are excluded, and every reported value remains traceable to a version-locked source table. No error bars are applicable. Please click here to view a larger version of this figure.

figure-results-2
Figure 2: Dataset profile and patient-level demographics. The summary boxes report 555 prescription visits, 396 unique patients, a patient-level age range of 13–94 years with a mean age of 47.11 ± 14.88 years, and 324 merged herb labels representing 9,339 herb-use occurrences. (A) Patient-level sex distribution among 396 unique patients, including female, male, and unknown categories. (B) Patient-level age distribution by sex among the same 396 patients. The summary boxes use the displayed visit-level or patient-level denominators, as applicable. Colors distinguish sex categories and summary elements and have no inferential meaning. Please click here to view a larger version of this figure.

figure-results-3
Figure 3: Disease spectra recovered from prescription visits. (A) Hierarchical Western disease groups and (B) hierarchical TCM disease groups derived from 555 eligible prescription visits. (C) Leading normalized Western diagnoses and (D) leading normalized TCM disease names from the same visit-level denominator. Ring segments and bars represent visit counts; colors distinguish diagnostic systems or categories and do not represent treatment effects. Retained TCM source labels are defined in Table 3: Wei pi (epigastric fullness or discomfort), Ji disease (accumulation-type source TCM label), Gan pi (a TCM disease label used for fatty-liver disorders), Gan zhuo (literally, liver fixity; a retained classical TCM disease label), Bi disease (impediment disease), Xiao ke (wasting-thirst), Xie xie (loose or watery stools), and Fu xie (diarrhea). These glosses do not convert the source labels into modern biomedical diagnoses. Please click here to view a larger version of this figure.

figure-results-4
Figure 4: Explicit syndrome-text spectrum and co-occurrence structure. (A) Visit-level frequencies of 14 normalized syndrome elements after within-visit deduplication, using 555 eligible prescription visits as the denominator. (B) Co-occurrence network of the 12 most frequent syndrome elements; node size represents visit frequency, and edge width and opacity represent pairwise co-occurrence counts. (C) Corresponding symmetric co-occurrence count matrix; diagonal cells represent individual visit frequencies, and off-diagonal cells represent pairwise co-occurrence counts. The audit strip reports 0 exact syndrome-text/indicator matches across 555 rows and a mean text-indicator Jaccard similarity of 0.1806; therefore, the legacy indicator columns were excluded from substantive reporting. In the matrix, Qi def. denotes qi deficiency, Qi stag. denotes qi stagnation, and Water ret. denotes water retention. The color palette is descriptive. No error bars are applicable. Please click here to view a larger version of this figure.

figure-results-5
Figure 5: High-frequency herbs and normalized materia-medica attribute profiles. (A) The twenty most frequent normalized herb labels among 555 eligible prescription visits. (B) Rank-frequency profile of 36 leading normalized herb labels. (C) Four-Qi distribution based on 9,115 property-coded herb-use occurrences; Four-Qi is single-valued for each coded occurrence. (D) Five-flavor and (E) meridian-tropism distributions based on the same property-coded denominator; both are multi-label attributes, and category counts are therefore nonexclusive. (F) Visit-level and property-coded analysis units and their reproducibility link. Bar lengths represent counts, and colors distinguish panels. No error bars are applicable. Please click here to view a larger version of this figure.

figure-results-6
Figure 6: History-derived symptom keyword spectrum and reproducible aggregation workflow. (A) Proportion of 555 eligible prescription visits containing at least one prespecified history keyword; 474 visits (85.41%) contained at least one keyword. (B) Fifteen most frequent exact history-derived symptom concepts. (C) Full 37-term rank-frequency distribution and cumulative share of keyword detections. (D) Version-locked workflow from the prespecified dictionary to aggregate outputs. Each concept was counted no more than once within a visit, whereas multiple concepts could occur in the same visit. Orange bars represent keyword counts, and workflow colors distinguish processing stages. No error bars are applicable. Please click here to view a larger version of this figure.

figure-results-7
Figure 7: Directed herb association rules and hierarchical clustering. (A) Directed bipartite network of all 15 single-herb association rules meeting the prespecified thresholds of support ≥ 0.30, confidence ≥ 0.70, and lift > 1.0. Edge width represents support, and edge color represents lift. (B) Rule-strength profile of the same 15 rules ranked by lift; point size represents support, and adjacent labels report confidence (conf.). (C) Average-linkage hierarchical clustering of the 20 highest-frequency herb-presence variables using Jaccard distance. All panels use prescription visits as the unit of analysis and describe co-recording patterns rather than efficacy, synergy, or a recommended fixed combination. No error bars are applicable. Please click here to view a larger version of this figure.

figure-results-8
Figure 8: Diagnosis-stratified aggregate herb patterns, clinical-pattern map, and release boundary. (A) Diagnosis-stratified heatmap showing visit-level prevalence of 16 normalized herb labels in fatty liver (n = 77), liver cirrhosis (n = 46), and chronic gastritis (n = 133). Each cell represents the percentage of eligible visits within the corresponding Western-diagnosis stratum, using a fixed 0%–100% color scale. (B) Record-derived clinical-pattern map linking verified record cues to normalized text groups and observed aggregate herb labels. Dashed, non-directional connectors distinguish this descriptive interpretation layer from computed prevalence; the panel does not constitute a treatment recommendation. (C) Reproducibility and release boundary distinguishing source-verified or reanalyzed components from items requiring caution or confirmation. Please click here to view a larger version of this figure.

MetricValueUnit/denominatorStatus
Prescription visits555555 prescription visitsverified from recovered XLSX
Unique patients396396 unique patientsverified from recovered XLSX
Patient-level age range13-94yearsverified
Patient-level age mean +/- SD47.11 +/- 14.88yearsverified
Visit-level age mean +/- SD48.22 +/- 15.82yearsverified
Patient-level sexFemale 228; male 167; unknown 1396 patientsverified
Visit-level sexFemale 327; male 227; unknown 1555 visitsverified
Merged herbs324distinct merged namesverified
Merged herb-use occurrences9339555 prescription visitsverified

Table 1: Dataset profile. Patient-level demographic characteristics and prescription-visit-level variables are reported using their respective denominators. Normalized herb totals are also provided.

Disease nameCountPercent of visits
Chronic gastritis13324
Fatty liver7713.9
Liver cirrhosis468.3
Chronic atrophic gastritis224
Abdominal pain224
Chronic hepatitis B203.6
Dyspepsia193.4
Reflux esophagitis81.4
Chronic superficial gastritis71.3
Chronic pancreatitis61.1
Ulcerative colitis61.1
Gallbladder polyp61.1
Post-cholecystectomy syndrome61.1
Intestinal dysfunction61.1
Gallstones61.1
Chronic hepatitis50.9
Fatigue50.9
Gastrointestinal dysfunction (uncertain original label)50.9
Chronic gastritis (uncertain original label)50.9
Autoimmune liver cirrhosis40.7

Table 2: Leading normalized Western diagnoses. Counts and percentages of leading normalized Western diagnoses recorded across 555 eligible prescription visits. Percentages use eligible prescription visits as the denominator.

TCM disease nameCountPercent of visits
Wei pi (epigastric fullness or discomfort)16329.4
Ji disease (accumulation-type source TCM label)8615.5
Gan pi (TCM disease label used for fatty-liver disorders)7713.9
Stomach pain6111
Hypochondriac pain325.8
Abdominal pain315.6
Gan zhuo (literally, liver fixity; a retained classical TCM disease label)285
Deficiency taxation173.1
Acid regurgitation91.6
Xie xie (loose or watery stools)61.1
Insomnia40.7
Dysentery40.7
Headache40.7
Bi disease (impediment disease)40.7
Menstrual disorder30.5
Chest impediment30.5
Sweating disorder30.5
Constipation30.5
Fu xie (diarrhea)20.4
Xiao ke (wasting-thirst)20.4

Table 3: Leading normalized TCM disease names. Counts and percentages of leading normalized TCM disease names recorded across 555 eligible prescription visits. Percentages use eligible prescription visits as the denominator. Retained source labels and explanatory glosses do not imply equivalence with modern biomedical diagnoses.

RankSyndrome elementPrescription visitsEligible visitsVisit percentage
1spleen36255565.23%
2liver30455554.77%
3qi deficiency29555553.15%
4stomach17155530.81%
5dampness15155527.21%
6qi stagnation10955519.64%
7heat9755517.48%
8water retention9455516.94%
9Blood8755515.68%
10blood stasis5855510.45%
11channels5755510.27%
12kidney435557.75%
13heart375556.67%
14yin deficiency345556.13%

Table 4: Leading normalized syndrome elements. Syndrome elements were extracted from explicit normalized syndrome text and deduplicated within each prescription visit. Counts use 555 eligible prescription visits as the denominator; percentages are nonexclusive because multiple syndrome elements can occur within a visit.

Merged herb labelCountPercent of visits
Bai Shao53195.7
Fu Ling52093.7
Dang Gui51091.9
Chi Shao46784.1
Stir-fried Bai Zhu36165
Vinegar-processed E Zhu30755.3
Tai Zi Shen28451.2
Dan Shen28250.8
Gan Cao27850.1
Huang Lian19935.9
Sheng Huang Qi17631.7
Mu Xiang17631.7
Chai Hu14425.9
Vinegar-processed Ji Nei Jin14225.6
Mu Dan Pi13724.7
Sheng Di Huang13123.6
Honey-fried Gan Cao12923.2
Fa Ban Xia11921.4
Chuan Xiong11320.4
Ginger-processed Hou Po10619.1
Yin Chen10418.7
Sheng Bai Zhu9717.5
Dou Kou9617.3
Sheng Long Gu9216.6
Zhe Bei Mu9016.2
Huang Qin8815.9
Dang Shen8515.3
Vinegar-processed Bie Jia8315
Gui Zhi8315
Sheng Mu Li8114.6
Jiao Zhi Zi7413.3
He Huan Hua6912.4
Wine-processed Huang Jing6611.9
Chen Pi6411.5
Ze Xie6311.4
Jin Qian Cao6111

Table 5: Leading normalized herb labels by prescription-visit frequency. Each normalized herb label was counted at most once per prescription visit. Counts and percentages use 555 eligible prescription visits as the denominator, and clinically meaningful processing forms are retained as separate labels.

RankHistory-derived keywordPrescription visitsEligible visitsVisit percentage
1Dry throat17855532.07%
2Fatigue15155527.21%
3Poor sleep13955525.05%
4Abdominal distension11555520.72%
5Acid reflux10055518.02%
6Bitter taste9955517.84%
7Heartburn9055516.22%
8Hiccup8755515.68%
9Belching8755515.68%
10Loose stool7755513.87%
11Diarrhea7755513.87%
12Epigastric distension7655513.69%
13Stomach pain7355513.15%
14Easy waking5955510.63%
15Dizziness545559.73%
16Vomiting525559.37%
17Nausea525559.37%
18Dull pain515559.19%
19Distending pain505559.01%
20Poor appetite445557.93%
21Abdominal pain435557.75%
22Chest oppression275554.86%
23Irritability265554.68%
24Headache235554.14%
25Borborygmus235554.14%
26Epigastric discomfort215553.78%
27Palpitations195553.42%
28Yellow urine185553.24%
29Foreign-body sensation155552.70%
30Stabbing pain145552.52%
31Constipation135552.34%
32Drowsiness125552.16%
33Pharyngeal foreign-body sensation95551.62%
34Excessive sweating85551.44%
35Hypochondriac pain55550.90%
36Tenesmus55550.90%
37Dry stool35550.54%

Table 6: Leading history-derived symptom keywords. Exact symptom concepts detected in de-identified history text. Each concept was counted no more than once within a prescription visit, whereas multiple concepts could occur within the same visit. Counts and percentages use 555 eligible prescription visits as the denominator.

AntecedentConsequentCo-occurring visitsSupportConfidenceLift
Chi ShaoFu Ling4380.78920.93791.0010
Fu LingChi Shao4380.78920.84231.0010
Stir-fried Bai ZhuFu Ling3500.63060.96951.0348
Stir-fried Bai ZhuBai Shao3480.62700.96401.0076
Stir-fried Bai ZhuChi Shao3200.57660.88641.0535
Tai Zi ShenFu Ling2730.49190.96131.0260
Tai Zi ShenBai Shao2720.49010.95771.0010
Gan CaoBai Shao2710.48830.97481.0189
Gan CaoChi Shao2620.47210.94241.1200
Tai Zi ShenChi Shao2420.43600.85211.0127
Gan CaoStir-fried Bai Zhu2030.36580.73021.1226
Huang LianFu Ling1870.33690.93971.0029
Huang LianChi Shao1800.32430.90451.0750
Sheng Huang QiDang Gui1680.30270.95451.0388
Mu XiangFu Ling1670.30090.94891.0127

Table 7: Directed herb association rules meeting the prespecified thresholds. Rules were generated from 555 binary prescription-visit rows using support ≥ 0.30, confidence ≥ 0.70, and lift > 1.0. Support uses 555 prescription visits as the denominator, and the rule direction is retained because confidence is directional.

Herb labelFatty liver (n)Fatty liver (N)Fatty liver (%)Liver cirrhosis (n)Liver cirrhosis (N)Liver cirrhosis (%)Chronic gastritis (n)Chronic gastritis (N)Chronic gastritis (%)
Bai Shao717792.21%444695.65%13013397.74%
Dang Gui727793.51%424691.30%11913389.47%
Dan Shen457758.44%444695.65%4313332.33%
Fu Ling707790.91%434693.48%12613394.74%
Vinegar-processed E Zhu577774.03%404686.96%6113345.86%
Vinegar-processed Bie Jia6777.79%394684.78%41333.01%
Stir-fried Bai Zhu577774.03%334671.74%7613357.14%
Chi Shao697789.61%324669.57%12113390.98%
Huang Lian237729.87%74615.22%5913344.36%
Mu Xiang207725.97%134628.26%4613334.59%
Yin Chen327741.56%134628.26%101337.52%
Sheng Huang Qi237729.87%284660.87%2113315.79%
Gan Cao457758.44%184639.13%7713357.89%
Chai Hu127715.58%3466.52%4813336.09%
Huang Qin167720.78%4468.70%2513318.80%
Ze Xie247731.17%94619.57%71335.26%

Table 8: Diagnosis-stratified herb-use prevalence. Numerators, denominators, and visit-level prevalence of 16 normalized herb labels in fatty liver (n = 77), liver cirrhosis (n = 46), and chronic gastritis (n = 133). Each herb label was counted no more than once per eligible prescription visit.

Supplementary Table 1: AI-assisted transparency and audit record. Documentation distinguishing the original-analysis record from the revision-stage consistency review conducted on 21 August 2026. The table records the tool/model/version, input scope, purpose, identified Four-Qi cardinality mismatch, human correction, use limitations, reviewer information, disposition, and prompt template. No raw patient-level rows were provided to the AI system. Please click here to download this file.

Supplementary Table 2: Source-audit data-quality outcomes. Recovered-field completeness, audit findings, analytical actions, and released states are reported for source clinical fields and legacy analytical structures. The table documents exclusion of unreconciled legacy syndrome indicators, correction of earlier herb totals, the absence of reconstructable pre-cleaning distribution vectors for cross-version stability estimation, and quantities that could not be reconstructed. Please click here to download this file.

Supplementary Table 3: Herb nomenclature mapping. Controlled herb display labels are linked to processing qualifiers, standard Chinese medicinal material (CMM) names, source terminology, botanical, zoological, or mineral sources, reference versions, evidence locators, and evidence boundaries. Processing-specific labels are retained where clinically meaningful. Please click here to download this file.

Supplementary Table 4: Threshold and repeat-visit sensitivity analyses. Association-rule threshold perturbations, one-visit-per-patient rule-set comparisons, Chi Shao > Fu Ling sensitivity metrics, and the 190-pair clustering-input distance comparison are reported. Please click here to download this file.

Supplementary File 1: Deterministic threshold and repeat-visit sensitivity script. Deterministic script used to reproduce aggregate threshold and repeat-visit sensitivity metrics from the version-locked source dataset. The file contains no patient-level data. Please click here to download this file.

Discussion

The central contribution of this protocol is a complete and inspectable pathway from de-identified TCM case records to aggregate research figures. The most critical steps are to freeze the cohort-identification rule before inspecting analytical results, distinguish unique patients from prescription visits, preserve every original-to-normalized term mapping, and verify each numerator and denominator before visualization. These controls transform data cleaning from an invisible preliminary activity into a documented component of the method. They also align the workflow with RECORD, STROBE, FAIR, and established EHR data-quality frameworks11,12,13,14,15. A result is considered ready for release only when its source field, analysis unit, dictionary version, calculation rule, denominator, table, figure, and reviewer decision can be traced in the evidence ledger.

The method is innovative because governance, terminology engineering, computation, visualization, and clinical review are integrated rather than treated as separate tasks. A single explicit syndrome matrix generates the frequency spectrum, co-occurrence network, and heatmap. A single prescription-visit-by-herb matrix supports frequency summaries, directed association rules, and hierarchical clustering19,20,21. A single diagnosis-stratified table supports both exact percentages and the final heatmap. The analytical workflow uses scripted, reproducible open-source tools36,37,38,39,40, allowing a correction to an upstream dictionary or cohort decision to propagate through all dependent outputs. Earlier clinical data warehouse work also demonstrates the importance of source-linked preprocessing and structured data reuse41,42. This architecture reduces manual transcription, maintains alignment between graphical encodings and source tables, and facilitates audit, teaching, and transfer to other research teams.

Three recent database-mining studies provide useful design principles, although their unit of analysis is a publication rather than a clinical visit. The rheumatoid arthritis and depression study combines a predefined search strategy with country, institution, author, journal, keyword, and timeline analyses26. The ulcerative colitis and Janus kinase inhibitor study uses coordinated collaboration, co-citation, clustering, and burst analyses to distinguish persistent themes from emerging research fronts27. The rheumatoid arthritis and fibromyalgia study extends bibliometric mapping with bioinformatics while maintaining a distinction between the two evidence layers28. The present protocol applies the same principle of coordinated but noninterchangeable analytical views. Transfer to gynecology, respiratory medicine, rheumatology, or another specialty requires reconstruction of specialty-specific fields, terminology dictionaries, inclusion and exclusion rules, outcomes, and clinical review procedures before the rule set is reused. Disease spectra, syndrome networks, symptom keywords, association rules, clustering, and diagnosis-stratified patterns all arise from version-locked sources, but each addresses a distinct question and retains its own denominator and interpretation boundary.

Human-supervised AI can provide an additional quality-control layer when restricted to approved de-identified excerpts or aggregate materials. Within this workflow, AI can compare terminology across files, flag inconsistencies between captions and tables, identify labels that extend beyond figure boundaries, verify that each displayed network node has a valid edge, and suggest clearer wording. AI does not determine cohort eligibility, invent terminology mappings, infer treatment effects, or replace source-data review. This division of responsibilities is consistent with the broader view that medical AI should augment human judgment rather than obscure accountability33. It also reflects the DECIDE-AI emphasis on transparent human factors, error handling, and accountable evaluation34. The purpose of each prompt, the reviewed output, the accepted correction, the reviewer, and the date should therefore be retained in the audit log whenever AI assistance is used.

Troubleshooting is particularly important when an output appears clinically plausible but fails source reconciliation. In such cases, downstream interpretation should be suspended until the discrepancy is traced to eligibility determination, deduplication, terminology mapping, field selection, denominator choice, or figure assembly. All dependent outputs should then be regenerated after correction. The worked example remains a single-center, retrospective, expert-outpatient dataset in which repeated visits are not independent of one another. The one-visit-per-patient sensitivity analysis reduced the retained rule set from 15 to 12 while preserving membership of all 20 top herbs. Across all 190 unordered pairs among these 20 shared variables, the Pearson correlation was 0.9944, Spearman rho was 0.9836, the mean absolute change in Jaccard distance was 0.0146, and the maximum change was 0.0528 (Supplementary Table 4). Accordingly, high visit-level support should not be interpreted as patient-level validation. The history-keyword spectrum reflects literal documentation rather than a validated symptom scale. Materia-medica summaries are occurrence-weighted rather than dose-weighted. Association rules describe co-recording rather than causation. Diagnosis-stratified heatmaps are descriptive and do not establish comparative effectiveness, disease specificity, or treatment necessity.

The protocol can be applied to research on the inheritance of expert clinical experience, the description of specialist outpatient cohorts, multicenter terminology harmonization, documentation-quality review, variable selection for prospective registries, and privacy-conscious preparation of aggregate collaboration datasets. Future work can incorporate prospective structured symptom collection, patient-clustered models, explicit dose variables, prespecified comparative hypotheses, external validation, versioned terminology services, privacy-preserving computation, and interactive evidence ledgers. Natural-language processing and AI-assisted review may further reduce repetitive curation when governed by version-locked rules and human verification. More broadly, the protocol enables dispersed clinical experience to be transformed into a transparent research object that can be inspected, reproduced, challenged, and extended. The method does not convert frequency into efficacy; rather, it converts undocumented analytical decisions into visible evidence, providing a foundation for more rigorous clinical questions and more durable real-world TCM research.

Disclosures

The authors declare no competing interests.

Acknowledgements

OpenAI Codex (gpt-5.6-luna and gpt-5.6-sol) was used exclusively during manuscript revision for aggregate-level cross-file consistency checks and wording suggestions. No patient-level records were provided to the AI system. All AI-assisted outputs were independently verified against the source materials by Tian Xia on 21 August 2026. The authors reviewed and approved the final content and take full responsibility for the manuscript. This project was supported by the National Key Research and Development Program of China, Key Special Project on Modernization of Traditional Chinese Medicine (No. 2025YFC3512800); the National Science and Technology Major Project of China (Project No. 2026ZD0554100; Subproject No. 2026ZD0554101); the Central High-level Traditional Chinese Medicine Hospital Clinical Research and Achievement Translation Capacity Improvement Project, Special Project for Inheriting the Academic Experience of Renowned Senior Traditional Chinese Medicine Experts (No. HLCMHPP2023016); and the Natural Science Foundation of Xinjiang Uygur Autonomous Region (No. 2025D01C213).

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
Aggregate reanalysis spreadsheetsAuthor-generated local filesNot applicableUsed for aggregate counts, tables, and figure generation.
De-identified outpatient case-record exportInstitutional electronic medical record systemNot applicableUsed only in the controlled institutional environment; patient-level records are not submission materials.
lxmllxml project6.1.1Software version used for namespace-preserving serialization during revision-stage regeneration; not a historical original-analysis version.
MatplotlibMatplotlib project (matplotlib.org)3.11.1Version observed in the revision-stage regeneration environment and Figure 3 SVG metadata; not a historical original-analysis version.
NetworkXNetworkX projectPyPI package: networkxPackage identifier recorded; exact revision-stage version was not verified in the inherited environment.
NumPyNumPy project2.3.5Software version used to regenerate the revised aggregate outputs; not a historical original-analysis version.
openpyxlopenpyxl project3.1.5Software version used to write the revised aggregate XLSX outputs; not a historical original-analysis version.
pandaspandas project3.0.1Software version used to regenerate the revised aggregate outputs; not a historical original-analysis version.
PyMuPDFPyMuPDF project1.28.2Software version used for revision-stage regeneration of revised figure outputs; not a historical original-analysis version.
PythonPython Software Foundation3.13.15Software version used to regenerate the revised aggregate outputs; not a historical original-analysis version.
SciPySciPy projectPyPI package: scipyPackage identifier recorded; exact revision-stage version was not verified in the inherited environment.

References

  1. Wang YH, et al. Study on knowledge discovery in traditional Chinese medical case records [in Chinese]. Zhong Xi Yi Jie He Xue Bao. 2007;5(4):368-372.
  2. Wang Y, et al. Automatic symptom name normalization in clinical records of traditional Chinese medicine. BMC Bioinformatics. 2010;11:40.
  3. Zhou L, et al. Natural language processing algorithms for normalizing expressions of synonymous symptoms in traditional Chinese medicine. Evid Based Complement Alternat Med. 2021;2021:6676607.
  4. Chu X, et al. Quantitative knowledge presentation models of traditional Chinese medicine (TCM): a review. Artif Intell Med. 2020;103:101810.
  5. Zhang T, et al. Constructing fine-grained entity recognition corpora based on clinical records of traditional Chinese medicine. BMC Med Inform Decis Mak. 2020;20:64.
  6. Wang Q, et al. A study of entity-linking methods for normalizing Chinese diagnosis and procedure terms to ICD codes. J Biomed Inform. 2020;105:103418.
  7. Zhang H, et al. Transformer- and generative adversarial network-based inpatient traditional Chinese medicine prescription recommendation: development study. JMIR Med Inform. 2022;10(5):e35239.
  8. Dan W, et al. SinoMedminer: an R package and Shiny application for mining and visualizing traditional Chinese medicine herbal formulas. Chin Med. 2025;20:80.
  9. Chinese Pharmacopoeia Commission. Pharmacopoeia of the People's Republic of China. 2025 ed. China Medical Science and Technology Press; Beijing; 2025.
  10. Editorial Committee of Zhonghua Bencao, State Administration of Traditional Chinese Medicine. Zhonghua Bencao. 10 vols. Shanghai Scientific and Technical Publishers; Shanghai; 1999. ISBN: 7532351068. Available from: National Diet Library record
  11. Benchimol EI, et al. The REporting of studies Conducted using Observational Routinely-collected health Data (RECORD) statement. PLoS Med. 2015;12(10):e1001885.
  12. von Elm E, et al. The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement: guidelines for reporting observational studies. Lancet. 2007;370(9596):1453-1457.
  13. Wilkinson MD, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:160018.
  14. Weiskopf NG, Weng C. Methods and dimensions of electronic health record data quality assessment: enabling reuse for clinical research. J Am Med Inform Assoc. 2013;20(1):144-151.
  15. Kahn MG, et al. A harmonized data quality assessment terminology and framework for the secondary use of electronic health record data. EGEMS (Wash DC). 2016;4(1):18.
  16. Kovačević A, Bašaragin B, Milošević N, Nenadić G. De-identification of clinical free text using natural language processing: a systematic review of current approaches. Artif Intell Med. 2024;151:102845.
  17. Meystre SM, et al. Text de-identification for privacy protection: a study of its impact on clinical text information content. J Biomed Inform. 2014;50:142-150.
  18. Neamatullah I, et al. Automated de-identification of free-text medical records. BMC Med Inform Decis Mak. 2008;8:32.
  19. Agrawal R, Srikant R. Fast algorithms for mining association rules in large databases. Presented at: 20th International Conference on Very Large Data Bases; Santiago, Chile; 1994. p. 487-499.
  20. Hahsler M, Grün B, Hornik K. arules: a computational environment for mining association rules and frequent item sets. J Stat Softw. 2005;14(15):1-25.
  21. Murtagh F, Contreras P. Algorithms for hierarchical clustering: an overview. WIREs Data Min Knowl Discov. 2012;2(1):86-97.
  22. van Eck NJ, Waltman L. Software survey: VOSviewer, a computer program for bibliometric mapping. Scientometrics. 2010;84(2):523-538.
  23. Chen C. CiteSpace II: detecting and visualizing emerging trends and transient patterns in scientific literature. J Am Soc Inf Sci Technol. 2006;57(3):359-377.
  24. Aria M, Cuccurullo C. bibliometrix: an R-tool for comprehensive science mapping analysis. J Informetr. 2017;11(4):959-975.
  25. Donthu N, et al. How to conduct a bibliometric analysis: an overview and guidelines. J Bus Res. 2021;133:285-296.
  26. Zhao Y, Chen GY, Fang M. Research trends of rheumatoid arthritis and depression from 2019 to 2023: a bibliometric analysis. J Multidiscip Healthc. 2024;17:4465-4474.
  27. Wang Y, et al. Research trends and hotspots in JAK inhibitors for ulcerative colitis: a bibliometric analysis from 2015 to 2024. J Multidiscip Healthc. 2025;18:6755-6771.
  28. Chen G, Yan Z, Wang Y, Tao Q. Rheumatoid arthritis and fibromyalgia syndrome: a bibliometric and bioinformatics perspective on comorbidity research. J Multidiscip Healthc. 2025;18:6811-6827.
  29. Wang SL, et al. Yao Naili's experience in applying the harmonizing liver and spleen method. Journal of Traditional Chinese Medicine. 2008;49(7):596-597.
  30. Wang L, et al. Analysis of Professor Naili Yao's experience in treating spleen-stomach diseases [in Chinese]. Lishizhen Med Mater Med Res. 2022;33(6):1436-1438.
  31. Sandve GK, Nekrutenko A, Taylor J, Hovig E. Ten simple rules for reproducible computational research. PLoS Comput Biol. 2013;9(10):e1003285.
  32. Munafò MR, et al. A manifesto for reproducible science. Nat Hum Behav. 2017;1:0021.
  33. Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med. 2019;25(1):44-56.
  34. Vasey B, et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat Med. 2022;28(5):924-933.
  35. Branch of Gastrointestinal Diseases, China Association of Chinese Medicine. Expert consensus on traditional Chinese medicine diagnosis and treatment of Gan pi (2023). Chinese Journal of Integrative Digestive Diseases. 2024;32(9):741-747.
  36. McKinney W. Data structures for statistical computing in Python [conference paper]. Presented at: 9th Python in Science Conference; Austin, TX; 2010. p. 56-61. doi:10.25080/Majora-92bf1922-00a.
  37. Harris CR, et al. Array programming with NumPy. Nature. 2020;585(7825):357-362.
  38. Virtanen P, et al. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat Methods. 2020;17(3):261-272.
  39. Hunter JD. Matplotlib: a 2D graphics environment. Comput Sci Eng. 2007;9(3):90-95.
  40. Hagberg AA, Schult DA, Swart PJ. Exploring network structure, dynamics, and function using NetworkX. Presented at: 7th Python in Science Conference; Pasadena, CA; 2008. p. 11-15. doi:10.25080/TCWV9851.
  41. Zhou X, et al. Development of traditional Chinese medicine clinical data warehouse for medical knowledge discovery and decision support. Artif Intell Med. 2010;48(2-3):139-152.
  42. Liu B, et al. Data processing and analysis in real-world traditional Chinese medicine clinical data: challenges and approaches. Stat Med. 2012;31(7):653-660.

Reprints and Permissions

Tags

Real-World Case RecordsSource-Audited ProtocolData NormalizationSyndrome Element SpectrumHerb FrequencyMateria Medica ProfilesAssociation RulesClinical Data Visualization