A subscription to JoVE is required to view this content. Sign in or start your free trial.

Research Article

Readability and Quality of Large Language Model Patient Education for Trigeminal Neuralgia: A Cross-Sectional Study

70 views

⸱

DOI:

10.3791/70833

⸱

August 21st, 2026

In This Article

Summary

Here, we present a protocol to systematically evaluate the readability, educational suitability, and quality of large language model–generated patient information on trigeminal neuralgia to benchmark models and guide patient-facing communication.

Abstract

Large language models (LLMs) are increasingly used for patient health education, yet the readability and educational quality of LLM-generated information on trigeminal neuralgia (TN) have been insufficiently evaluated. This cross-sectional benchmarking study compared TN educational content generated by five publicly available LLMs (Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and GPT-5). Twenty frequently asked TN questions covering basic disease knowledge, etiology/risk factors, diagnosis, treatment, and prevention/rehabilitation were presented to each model using standardized prompts. Readability was assessed using seven established indices, educational suitability using the Patient Education Materials Assessment Tool for Understandability and Actionability (PEMAT), and overall information quality using the Global Quality Score (GQS). Two clinical experts independently evaluated all responses, with disagreements resolved by a senior adjudicator. Statistical analyses compared model performance, thematic differences, and correlations among the evaluation metrics. Significant differences were observed among the models for readability, PEMAT, and GQS scores. GPT-5 generated the most linguistically complex responses but achieved the highest ratings for educational suitability and information quality. In contrast, Wenxin Yiyan produced the most readable text but generally scored lower on PEMAT and GQS. Content category influenced readability, with prevention/rehabilitation and etiology/risk-factor topics being more difficult to read, whereas PEMAT and GQS remained relatively consistent across themes. Readability indices showed strong internal consistency and weak-to-moderate positive correlations with PEMAT and GQS, while PEMAT and GQS demonstrated a moderate positive correlation. These findings suggest that model selection influences expert-rated educational suitability and overall information quality, whereas topic complexity primarily affects readability. Because patient comprehension, satisfaction, trust, health outcomes, factual accuracy, and clinical safety were not evaluated, these results should be interpreted as an expert-rated benchmarking analysis rather than evidence of clinical readiness.

Introduction

Trigeminal neuralgia (TN) is a neuropathic pain syndrome characterized by recurrent episodes of unilateral, brief, and extremely severe facial pain commonly described as electric shock-like or stabbing. According to the International Classification of Headache Disorders, 3rd edition (ICHD-3), pain is typically localized to the distribution of one or more branches of the trigeminal nerve, is often intolerable in intensity, and can be triggered by innocuous stimuli such as light touch, facial washing, tooth brushing, speaking, or chewing. Attacks are marked by sudden onset and abrupt cessation, with durations ranging from fractions of a second to several minutes1,2. Although TN is not usually life-threatening, its severe and unpredictable pain substantially disrupts essential daily activities, including eating, speech, sleep, and emotional regulation, resulting in a considerable psychosocial burden and impaired quality of life. Evidence from systematic reviews indicates that individuals with TN face a significantly elevated risk of depression, anxiety, and sleep disturbances compared to the general population3,4. In large cohort studies, approximately 35–55% of patients with TN exhibit clinically meaningful symptoms of anxiety or depression, and pain catastrophizing is highly prevalent, suggesting a bidirectional interaction between chronic severe pain, insomnia, and affective disorders5,6. Epidemiological estimates indicate that the prevalence of TN in the general population ranges from approximately 0.03% to 0.3%, with a higher occurrence among women and older adults and notable variations across geographic regions, ethnic groups, and age strata7,8. In highly populated countries, such as China and India, rapid population aging has been accompanied by a sustained increase in the number of individuals affected by TN, thereby imposing a growing burden on healthcare systems and underscoring the need for more effective, scalable, and data-informed approaches to disease assessment and management8,9.

TN can be classified into classical, secondary, and idiopathic subtypes, with increasing evidence demonstrating substantial differences in the etiological mechanisms and therapeutic strategies among these categories1,10. Current studies suggest that neurovascular conflict-related structural changes at the trigeminal nerve root entry zone, peripheral nerve hyperexcitability, and altered central pain modulation jointly contribute to the pathogenesis of the disease11,12. Clinical guidelines recommend carbamazepine and oxcarbazepine as first-line treatments, whereas surgical or interventional approaches are considered when pharmacological therapy is ineffective or poorly tolerated10. Despite these therapeutic advances, TN is characterized by a relapsing–remitting course, and long-term pain recurrence remains common, making permanent remission difficult to achieve13. Consequently, long-term follow-up and structured chronic disease management are essential for monitoring the risk of recurrence, treatment response, and complications. Longitudinal cohort studies indicate that under standardized management strategies (including pharmacological treatment, surgical intervention when indicated, and comprehensive follow-up), approximately half of patients receiving conservative management can achieve a ≥50% reduction in overall pain burden within two years, without inevitable disease progression14. These findings underscore the importance of sustained patient engagement and effective health education for long-term disease management. Evidence suggests that patients with a better understanding of disease etiology, pathophysiology, and treatment options are more likely to actively participate in and adhere to therapeutic regimens, leading to improved outcomes15. However, traditional health education approaches are often limited in reach and scalability. With the widespread adoption of the internet, particularly online question-and-answer platforms and social media, patients increasingly rely on digital sources to obtain medical information16,17. Nevertheless, substantial variability in individual health literacy and the ability to access, process, and understand basic health information for informed decision-making pose major barriers to effective patient education18. Moreover, the uneven quality of online health information, including inaccurate or misleading content, may negatively influence patients’ medical decisions and psychological well-being19.

Artificial intelligence (AI) technologies, particularly the rapid development of large language models (LLMs) in recent years, have created new opportunities for medical science communication and health education20. Trained on large-scale textual corpora, these models can generate structured and logically coherent responses through natural language interaction21. Representative international systems, such as ChatGPT, have demonstrated promising potential for medical question answering, patient consultation, and medical education22,23. Several functionally similar LLMs have emerged in China, with continuously expanding user coverage and application scenarios24. Despite these advances, the application of LLMs in medical science popularization still faces substantial challenges, including the reliability of training data, update frequency, the scientific accuracy of generated responses, and a lack of clinical validation25. In addition, differences among models in language style, readability, logical structure, and citation accuracy may directly influence patients’ comprehension and trust in health-related information26. To date, systematic evaluations of LLM performance in TN-related health education have been limited.

Therefore, this study aimed to systematically evaluate and compare the performance of four widely used Chinese LLMs (Doubao, DeepSeek, Wenxin Yiyan, and Tongyi Qianwen) and a representative international LLM (ChatGPT) in answering patient education questions related to TN. Using a cross-sectional design, we constructed a TN question set based on clinical guidelines and common patient concerns, and assessed the responses generated by five publicly available LLMs (Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and GPT-5). Readability was evaluated using multiple metrics, and understandability and actionability were assessed using the Patient Education Materials Assessment Tool (PEMAT). Overall information quality was evaluated using the Global Quality Score (GQS). In addition, we analyzed how different health education topics influence these evaluation metrics and their interrelationships, aiming to provide evidence-based guidance for selecting and optimizing LLMs in clinical health communication.

Access restricted. Please log in or start a trial to view this content.

Protocol

This study analyzed only AI-generated text and did not involve human participants, animals, biological specimens, medical records, identifiable personal information, or intervention in clinical care. Therefore, ethics review and formal informed consent were not applicable.

Study design

This cross-sectional study evaluated the readability, quality, and educational suitability of responses generated by five LLMs in response to frequently asked questions about TN.

Identification of commonly asked TN-related questions

On November 25, 2025, two clinical experts with extensive experience in pain medicine independently reviewed the current clinical guidelines for trigeminal neuralgia (TN), including the International Classification of Headache Disorders, 3rd edition (ICHD-3), the European Academy of Neurology guideline, and the American Academy of Neurology/European Federation of Neurological Societies guideline1,10,11. Following a consensus discussion, they compiled a list of 20 TN-related questions commonly encountered in clinical practice. The number of questions was determined a priori to provide balanced coverage of the five predefined thematic domains, with four questions selected for each domain, while keeping the total number of model-generated responses within a range feasible for manual expert rating and verification. To improve reproducibility, the selection process followed a predefined domain-based framework: the experts first identified candidate questions within each domain, independently reviewed them for clinical relevance, patient-oriented wording, and avoidance of redundancy, and then reached consensus on the final four questions per domain. The final list of questions is provided in Table 1 and Supplementary File 1. The five predefined thematic domains were basic disease knowledge, etiology and risk factors, diagnosis, treatment, and prevention and rehabilitation. Because the question set was not derived from patient search logs, online search queries, or formal patient interviews, the potential for selection bias is acknowledged as a study limitation.

AI models and data collection procedure

On November 25, 2025, the finalized set of 20 TN-related questions was submitted to five publicly accessible large language model (LLM) platforms: Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and GPT-5/ChatGPT. All models were accessed through their standard public web-based chat interfaces; application programming interface (API) access was not used. For each platform, the default publicly available model available on the date of data collection was used without modification of any generation parameters. Because the public web interfaces did not display backend model snapshot identifiers, hidden system prompts, temperature, top-p, maximum output tokens, or other generation parameters, these items were recorded as "not displayed in the public web interface".

The prompt and response languages were English for all models. Each standardized prompt consisted solely of the verbatim English question without any additional model-specific instructions. Each question was submitted individually in a new, independent chat session with no previous conversation history, uploaded files, plug-ins, customized instructions, or follow-up prompts. The complete list of standardized prompts and the corresponding raw model outputs is provided in Supplementary File 1. Model metadata and data collection settings are summarized in Supplementary Table 1.

Readability assessment

The readability of the LLM-generated responses was evaluated using multiple established readability formulas available through an online readability assessment platform (http://readabilityformulas.com/). Given the absence of a universally accepted gold standard for readability assessment and consistent with previous studies, a comprehensive set of widely used readability indices was employed23,27. Seven indices were selected to capture complementary aspects of readability, including sentence length, word length, syllable burden, character-based complexity, reading ease, and estimated grade level, thereby reducing reliance on any single formula. Specifically, the Coleman-Liau Index (CL), Linsear Write Formula (LW), Automated Readability Index (ARI), Simple Measure of Gobbledygook (SMOG), Gunning Fog Index (GFOG), Flesch Reading Ease Score (FRES), and Flesch-Kincaid Grade Level (FKGL) were calculated. These indices estimate reading difficulty or grade level and provide quantitative measures of text readability (Table 2).

Evaluation of educational suitability and information quality

Educational suitability was assessed using the Patient Education Materials Assessment Tool (PEMAT), which comprises two domains: understandability and actionability28. The instrument includes 24 items, with 16 evaluating understandability and eight evaluating actionability. Each item was scored dichotomously (0 = criterion not met; 1 = criterion met), yielding a total score ranging from 0 to 24, with higher scores indicating greater suitability for patient education. Overall information quality was evaluated using the Global Quality Score (GQS)27, a widely used five-point Likert scale for assessing the quality of medical information (Table 3).

On November 25, 2025, two clinical experts with more than 3 years of experience in TN management evaluated all AI-generated responses using both assessment instruments. Before scoring, all AI-generated responses were anonymized with coded response identifiers, and reviewers were blinded to the LLM platform during PEMAT and GQS assessments. Disagreements were resolved by a third senior expert, who adjudicated the final scores. Inter-rater reliability was assessed using Cohen's kappa coefficient, with values >0.75 indicating excellent agreement, 0.40–0.75 indicating acceptable agreement, and <0.40 indicating poor agreement. Cohen's kappa values for both the PEMAT and GQS exceeded 0.75, demonstrating excellent inter-rater reliability.

Statistical analysis

All statistical analyses were performed using R software (version 4.4.3). Data normality was assessed using the Shapiro-Wilk test and Q-Q plots. Normally distributed variables were summarized as mean ± standard deviation and compared using one-way analysis of variance (ANOVA), followed by Tukey's post hoc test when appropriate. Non-normally distributed variables were summarized as medians and interquartile ranges and compared using the Kruskal-Wallis test, followed by Dunn's post hoc test with a Bonferroni correction for pairwise comparisons. Correlations among readability indices, PEMAT scores, and GQS values were evaluated using Spearman's rank correlation coefficient, with the Benjamini-Hochberg correction applied for multiple-testing. A two-sided adjusted P value < 0.05 was considered statistically significant.

Access restricted. Please log in or start a trial to view this content.

Results

This study systematically evaluated the effects of large language models (LLMs) and different categories of health education content on the readability and quality of generated texts. First, we compared the performance of five LLMs—Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and GPT-5—with respect to patient education suitability, overall information quality, and multiple readability metrics, including the PEMAT score, GQS, and established readability indices (ARI, FRES, GFOG, FKGL, CL, SMOG, and LW). Second, we exam...

Access restricted. Please log in or start a trial to view this content.

Discussion

This cross-sectional benchmarking study systematically compared the readability, educational suitability, and overall information quality of trigeminal neuralgia (TN) patient education materials generated by five publicly available large language models (LLMs). Four principal findings emerged. First, readability differed substantially across the models, with GPT-5 generating the most linguistically complex responses and Wenxin Yiyan producing the most readable content. Second, readability and expert-rated information qua...

Access restricted. Please log in or start a trial to view this content.

Disclosures

The authors declare that they have no competing interests.

Acknowledgements

This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. The authors thank the clinical colleagues who provided feedback on the trigeminal neuralgia question set and assisted in refining the evaluation framework. We also thank all contributors who supported manuscript preparation and revision.

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
DeepSeek chat platformDeepSeekhttps://chat.deepseek.com/Publicly accessible large language model interface used to generate responses
Doubao chat platformDoubaohttps://www.doubao.com/chat/Publicly accessible large language model interface used to generate responses
GPT 5 OpenAIhttps://chat.openai.com/Publicly accessible large language model interface used to generate responses
R software, version 4.4.3R Foundation for Statistical ComputingNot applicableStatistical analysis and visualization
Readability Formulas online readability calculatorReadability FormulasNot applicableOnline tool used to compute readability indices
Tongyi Qianwen chat platformTongyi Qianwenhttps://www.tongyi.com/Publicly accessible large language model interface used to generate responses
Wenxin Yiyan chat platformWenxin Yiyanhttps://yiyan.baidu.com/Publicly accessible large language model interface used to generate responses

References

  1. Headache Classification Committee of the International Headache Society (IHS). The International Classification of Headache Disorders, 3rd edition. Cephalalgia. 2018;38(1):1-211.
  2. Ashina S, et al. Trigeminal neuralgia. Nat Rev Dis Primers. 2024;10(1):39.
  3. Donertas-Ayaz B, Caudle RM. Locus coeruleus-noradrenergic modulation of trigeminal pain: Implications for trigeminal neuralgia and psychiatric comorbidities. Neurobiol Pain. 2023;13:100124.
  4. Martinelli R, et al. Psychological assessment in patients affected by trigeminal neuralgia: A systematic review. Neurosurg Rev. 2025;48(1):414.
  5. Melek LN, et al. Comparison of the neuropathic pain symptoms and psychosocial impacts of trigeminal neuralgia and painful posttraumatic trigeminal neuropathy. J Oral Facial Pain Headache. 2019;33(1):77-88.
  6. Zakrzewska JM, et al. Evaluating the impact of trigeminal neuralgia. Pain. 2017;158(6):1166-1174.
  7. Gambeta E, Chichorro JG, Zamponi GW. Trigeminal neuralgia: An overview from pathophysiology to pharmacological treatments. Mol Pain. 2020;16:1744806920901890.
  8. De Toledo IP, et al. Prevalence of trigeminal neuralgia: A systematic review. J Am Dent Assoc. 2016;147(7):570-576.e2.
  9. Svedung Wettervik T, et al. Incidence of trigeminal neuralgia: A population-based study in central Sweden. Eur J Pain. 2023;27(5):580-587.
  10. Bendtsen L, et al. European Academy of Neurology guideline on trigeminal neuralgia. Eur J Neurol. 2019;26(6):831-849.
  11. Cruccu G, et al. Trigeminal neuralgia: New classification and diagnostic grading for practice and research. Neurology. 2016;87(2):220-228.
  12. Araya EI, et al. Trigeminal neuralgia: Basic and clinical aspects. Curr Neuropharmacol. 2020;18(2):109-119.
  13. Chong MS, Bahra A, Zakrzewska JM. Guidelines for the management of trigeminal neuralgia. Cleve Clin J Med. 2023;90(6):355-362.
  14. Heinskou TB, et al. Favourable prognosis of trigeminal neuralgia when enrolled in a multidisciplinary management program: A two-year prospective real-life study. J Headache Pain. 2019;20(1):23.
  15. Liu S, et al. Comparative evaluation of ChatGPT and Gemini in brain-computer interface patient education: A multidimensional analysis of reliability, accuracy, comprehensibility, and readability. Int J Med Inform. 2026;206:106164.
  16. Forgie EME, et al. Social media and the transformation of the physician-patient relationship: Viewpoint. J Med Internet Res. 2021;23(12):e25230.
  17. Gantenbein L, Navarini AA, Maul LV, Brandt O, Mueller SM. Internet and social media use in dermatology patients: Search behavior and impact on the patient-physician relationship. Dermatol Ther. 2020;33(6):e14098.
  18. Youssef Y, et al. Social media and internet use among orthopedic patients in Germany: A multicenter survey. Front Digit Health. 2025;7:1486296.
  19. De Martino I, et al. Social media for patients: Benefits and drawbacks. Curr Rev Musculoskelet Med. 2017;10(1):141-145.
  20. Johnson KB, et al. Precision medicine, AI, and the future of personalized health care. Clin Transl Sci. 2021;14(1):86-93.
  21. Hirani R, et al. Artificial intelligence and healthcare: A journey through history, present innovations, and future possibilities. Life (Basel). 2024;14(5):557.
  22. Karatas G, Kirik F, Karatas ME, Cakir A, Ozdemir H. Comparative performance of ChatGPT o3-mini-high and DeepSeek-R1 in ophthalmology: An evaluation of diagnostic reasoning and case-based problem solving. J Fr Ophtalmol. 2025;49(1):104712.
  23. Hanci V, et al. Assessment of readability, reliability, and quality of ChatGPT, Bard, Gemini, Copilot, and Perplexity responses on palliative care. Medicine (Baltimore). 2024;103(33):e39305.
  24. Liu C, et al. Potential to perpetuate social biases in health care by Chinese large language models: A model evaluation study. Int J Equity Health. 2025;24(1):206.
  25. Haupt CE, Marks M. AI-generated medical advice—GPT and beyond. JAMA. 2023;329(16):1349-1350.
  26. Aydin S, Karabacak M, Vlachos V, Margetis K. Large language models in patient education: A scoping review of applications in medicine. Front Med (Lausanne). 2024;11:1477898.
  27. Kara M, et al. Evaluating the readability, quality, and reliability of responses generated by ChatGPT, Gemini, and Perplexity on the most commonly asked questions about ankylosing spondylitis. PLoS One. 2025;20(6):e0326351.
  28. Perez Rivera LR, et al. Evaluating the quality and reliability of large language models for plastic surgery patient education: A comparative analysis of ChatGPT and OpenEvidence. Aesthet Surg J. 2025. doi:10.1093/asj/sjaf249.
  29. Wilhelm TI, Roos J, Kaczmarczyk R. Large language models for therapy recommendations across three clinical specialties: Comparative study. J Med Internet Res. 2023;25:e49324.
  30. Will J, et al. Enhancing the readability of online patient education materials using large language models: Cross-sectional study. J Med Internet Res. 2025;27:e69955.
  31. Tukur Jido J, Al-Wizni A, Aung SL. Readability of AI-generated patient information leaflets on Alzheimer's disease, vascular dementia, and delirium. Cureus. 2025;17(6):e85463.
  32. Luo Z, Lin C, Kim TH, Shin YS, Ahn ST. Quality and readability analysis of artificial intelligence-generated medical information related to prostate cancer: A cross-sectional study of ChatGPT and DeepSeek. World J Mens Health. 2025. doi:10.5534/wjmh.250144.
  33. Joseph P, Silva NA, Nanda A, Gupta G. Evaluating the readability of online patient education materials for trigeminal neuralgia. World Neurosurg. 2020;144:e934-e938.
  34. Dong C, et al. Comparative evaluation of large language models in delivering guideline-compliant recommendations for topical NSAID use in musculoskeletal pain: A multidimensional analysis. Clin Rheumatol. 2025;44(11):4703-4710.
  35. Ozduran E, Buyukcoban S. Evaluating the readability, quality, and reliability of online patient education materials on post-COVID pain. PeerJ. 2022;10:e13686.
  36. Wang LW, Miller MJ, Schmitt MR, Wen FK. Assessing readability formula differences with written health information materials: Application, results, and recommendations. Res Social Adm Pharm. 2013;9(5):503-516.
  37. Singh SP, et al. Comprehension profile of patient education materials in endocrine care. Kans J Med. 2022;15(2):247-252.
  38. Marder RS, et al. ChatGPT-3.5 and ChatGPT-4.0 do not reliably create readable patient education materials for common orthopaedic upper- and lower-extremity conditions. Arthrosc Sports Med Rehabil. 2025;7(1):101027.
  39. Nasra M, et al. Can artificial intelligence improve patient educational material readability? A systematic review and narrative synthesis. Intern Med J. 2025;55(1):20-34.
  40. Bhatt C, et al. Evaluating readability, understandability, and actionability of online printable patient education materials for cholesterol management: A systematic review. J Am Heart Assoc. 2024;13(8):e030140.
  41. Avra TD, Le M, Hernandez S, Thure K, Ulloa JG. Readability assessment of online peripheral artery disease education materials. J Vasc Surg. 2022;76(6):1728-1732.
  42. Rooney MK, et al. Readability of patient education materials from high-impact medical journals: A 20-year analysis. J Patient Exp. 2021;8:2374373521998847.
  43. Ahmadzadeh K, et al. Patient education information material assessment criteria: A scoping review. Health Info Libr J. 2023;40(1):3-28.
  44. Duan L, Yao Z, Li X, Wu Y, Sheng D. Comparing large language models and human doctors in symptom-driven online medical consultations: A case study on trigeminal neuralgia. Digit Health. 2025;11:20552076251388140.
  45. Werner C, Harden J, Lawton J. Pathways to a diagnosis of trigeminal neuralgia: A qualitative study of patients' experiences. BMC Prim Care. 2025;26(1):65.

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Tags

Large Language ModelsReadability AssessmentEducational QualityPEMAT EvaluationGlobal Quality ScoreHealth Information QualityModel BenchmarkingPatient Comprehension