Here, we present a protocol to systematically evaluate the readability, educational suitability, and quality of large language model–generated patient information on trigeminal neuralgia to benchmark models and guide patient-facing communication.
A subscription to JoVE is required to view this content. Sign in or start your free trial.
Research Article
Here, we present a protocol to systematically evaluate the readability, educational suitability, and quality of large language model–generated patient information on trigeminal neuralgia to benchmark models and guide patient-facing communication.
Large language models (LLMs) are increasingly used for patient health education, yet the readability and educational quality of LLM-generated information on trigeminal neuralgia (TN) have been insufficiently evaluated. This cross-sectional benchmarking study compared TN educational content generated by five publicly available LLMs (Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and GPT-5). Twenty frequently asked TN questions covering basic disease knowledge, etiology/risk factors, diagnosis, treatment, and prevention/rehabilitation were presented to each model using standardized prompts. Readability was assessed using seven established indices, educational suitability using the Patient Education Materials Assessment Tool for Understandability and Actionability (PEMAT), and overall information quality using the Global Quality Score (GQS). Two clinical experts independently evaluated all responses, with disagreements resolved by a senior adjudicator. Statistical analyses compared model performance, thematic differences, and correlations among the evaluation metrics. Significant differences were observed among the models for readability, PEMAT, and GQS scores. GPT-5 generated the most linguistically complex responses but achieved the highest ratings for educational suitability and information quality. In contrast, Wenxin Yiyan produced the most readable text but generally scored lower on PEMAT and GQS. Content category influenced readability, with prevention/rehabilitation and etiology/risk-factor topics being more difficult to read, whereas PEMAT and GQS remained relatively consistent across themes. Readability indices showed strong internal consistency and weak-to-moderate positive correlations with PEMAT and GQS, while PEMAT and GQS demonstrated a moderate positive correlation. These findings suggest that model selection influences expert-rated educational suitability and overall information quality, whereas topic complexity primarily affects readability. Because patient comprehension, satisfaction, trust, health outcomes, factual accuracy, and clinical safety were not evaluated, these results should be interpreted as an expert-rated benchmarking analysis rather than evidence of clinical readiness.
Trigeminal neuralgia (TN) is a neuropathic pain syndrome characterized by recurrent episodes of unilateral, brief, and extremely severe facial pain commonly described as electric shock-like or stabbing. According to the International Classification of Headache Disorders, 3rd edition (ICHD-3), pain is typically localized to the distribution of one or more branches of the trigeminal nerve, is often intolerable in intensity, and can be triggered by innocuous stimuli such as light touch, facial washing, tooth brushing, speaking, or chewing. Attacks are marked by sudden onset and abrupt cessation, with durations ranging from fractions of a second to several minutes1,2. Although TN is not usually life-threatening, its severe and unpredictable pain substantially disrupts essential daily activities, including eating, speech, sleep, and emotional regulation, resulting in a considerable psychosocial burden and impaired quality of life. Evidence from systematic reviews indicates that individuals with TN face a significantly elevated risk of depression, anxiety, and sleep disturbances compared to the general population3,4. In large cohort studies, approximately 35–55% of patients with TN exhibit clinically meaningful symptoms of anxiety or depression, and pain catastrophizing is highly prevalent, suggesting a bidirectional interaction between chronic severe pain, insomnia, and affective disorders5,6. Epidemiological estimates indicate that the prevalence of TN in the general population ranges from approximately 0.03% to 0.3%, with a higher occurrence among women and older adults and notable variations across geographic regions, ethnic groups, and age strata7,8. In highly populated countries, such as China and India, rapid population aging has been accompanied by a sustained increase in the number of individuals affected by TN, thereby imposing a growing burden on healthcare systems and underscoring the need for more effective, scalable, and data-informed approaches to disease assessment and management8,9.
TN can be classified into classical, secondary, and idiopathic subtypes, with increasing evidence demonstrating substantial differences in the etiological mechanisms and therapeutic strategies among these categories1,10. Current studies suggest that neurovascular conflict-related structural changes at the trigeminal nerve root entry zone, peripheral nerve hyperexcitability, and altered central pain modulation jointly contribute to the pathogenesis of the disease11,12. Clinical guidelines recommend carbamazepine and oxcarbazepine as first-line treatments, whereas surgical or interventional approaches are considered when pharmacological therapy is ineffective or poorly tolerated10. Despite these therapeutic advances, TN is characterized by a relapsing–remitting course, and long-term pain recurrence remains common, making permanent remission difficult to achieve13. Consequently, long-term follow-up and structured chronic disease management are essential for monitoring the risk of recurrence, treatment response, and complications. Longitudinal cohort studies indicate that under standardized management strategies (including pharmacological treatment, surgical intervention when indicated, and comprehensive follow-up), approximately half of patients receiving conservative management can achieve a ≥50% reduction in overall pain burden within two years, without inevitable disease progression14. These findings underscore the importance of sustained patient engagement and effective health education for long-term disease management. Evidence suggests that patients with a better understanding of disease etiology, pathophysiology, and treatment options are more likely to actively participate in and adhere to therapeutic regimens, leading to improved outcomes15. However, traditional health education approaches are often limited in reach and scalability. With the widespread adoption of the internet, particularly online question-and-answer platforms and social media, patients increasingly rely on digital sources to obtain medical information16,17. Nevertheless, substantial variability in individual health literacy and the ability to access, process, and understand basic health information for informed decision-making pose major barriers to effective patient education18. Moreover, the uneven quality of online health information, including inaccurate or misleading content, may negatively influence patients’ medical decisions and psychological well-being19.
Artificial intelligence (AI) technologies, particularly the rapid development of large language models (LLMs) in recent years, have created new opportunities for medical science communication and health education20. Trained on large-scale textual corpora, these models can generate structured and logically coherent responses through natural language interaction21. Representative international systems, such as ChatGPT, have demonstrated promising potential for medical question answering, patient consultation, and medical education22,23. Several functionally similar LLMs have emerged in China, with continuously expanding user coverage and application scenarios24. Despite these advances, the application of LLMs in medical science popularization still faces substantial challenges, including the reliability of training data, update frequency, the scientific accuracy of generated responses, and a lack of clinical validation25. In addition, differences among models in language style, readability, logical structure, and citation accuracy may directly influence patients’ comprehension and trust in health-related information26. To date, systematic evaluations of LLM performance in TN-related health education have been limited.
Therefore, this study aimed to systematically evaluate and compare the performance of four widely used Chinese LLMs (Doubao, DeepSeek, Wenxin Yiyan, and Tongyi Qianwen) and a representative international LLM (ChatGPT) in answering patient education questions related to TN. Using a cross-sectional design, we constructed a TN question set based on clinical guidelines and common patient concerns, and assessed the responses generated by five publicly available LLMs (Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and GPT-5). Readability was evaluated using multiple metrics, and understandability and actionability were assessed using the Patient Education Materials Assessment Tool (PEMAT). Overall information quality was evaluated using the Global Quality Score (GQS). In addition, we analyzed how different health education topics influence these evaluation metrics and their interrelationships, aiming to provide evidence-based guidance for selecting and optimizing LLMs in clinical health communication.
Access restricted. Please log in or start a trial to view this content.
This study analyzed only AI-generated text and did not involve human participants, animals, biological specimens, medical records, identifiable personal information, or intervention in clinical care. Therefore, ethics review and formal informed consent were not applicable.
Study design
This cross-sectional study evaluated the readability, quality, and educational suitability of responses generated by five LLMs in response to frequently asked questions about TN.
Identification of commonly asked TN-related questions
On November 25, 2025, two clinical experts with extensive experience in pain medicine independently reviewed the current clinical guidelines for trigeminal neuralgia (TN), including the International Classification of Headache Disorders, 3rd edition (ICHD-3), the European Academy of Neurology guideline, and the American Academy of Neurology/European Federation of Neurological Societies guideline1,10,11. Following a consensus discussion, they compiled a list of 20 TN-related questions commonly encountered in clinical practice. The number of questions was determined a priori to provide balanced coverage of the five predefined thematic domains, with four questions selected for each domain, while keeping the total number of model-generated responses within a range feasible for manual expert rating and verification. To improve reproducibility, the selection process followed a predefined domain-based framework: the experts first identified candidate questions within each domain, independently reviewed them for clinical relevance, patient-oriented wording, and avoidance of redundancy, and then reached consensus on the final four questions per domain. The final list of questions is provided in Table 1 and Supplementary File 1. The five predefined thematic domains were basic disease knowledge, etiology and risk factors, diagnosis, treatment, and prevention and rehabilitation. Because the question set was not derived from patient search logs, online search queries, or formal patient interviews, the potential for selection bias is acknowledged as a study limitation.
AI models and data collection procedure
On November 25, 2025, the finalized set of 20 TN-related questions was submitted to five publicly accessible large language model (LLM) platforms: Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and GPT-5/ChatGPT. All models were accessed through their standard public web-based chat interfaces; application programming interface (API) access was not used. For each platform, the default publicly available model available on the date of data collection was used without modification of any generation parameters. Because the public web interfaces did not display backend model snapshot identifiers, hidden system prompts, temperature, top-p, maximum output tokens, or other generation parameters, these items were recorded as "not displayed in the public web interface".
The prompt and response languages were English for all models. Each standardized prompt consisted solely of the verbatim English question without any additional model-specific instructions. Each question was submitted individually in a new, independent chat session with no previous conversation history, uploaded files, plug-ins, customized instructions, or follow-up prompts. The complete list of standardized prompts and the corresponding raw model outputs is provided in Supplementary File 1. Model metadata and data collection settings are summarized in Supplementary Table 1.
Readability assessment
The readability of the LLM-generated responses was evaluated using multiple established readability formulas available through an online readability assessment platform (http://readabilityformulas.com/). Given the absence of a universally accepted gold standard for readability assessment and consistent with previous studies, a comprehensive set of widely used readability indices was employed23,27. Seven indices were selected to capture complementary aspects of readability, including sentence length, word length, syllable burden, character-based complexity, reading ease, and estimated grade level, thereby reducing reliance on any single formula. Specifically, the Coleman-Liau Index (CL), Linsear Write Formula (LW), Automated Readability Index (ARI), Simple Measure of Gobbledygook (SMOG), Gunning Fog Index (GFOG), Flesch Reading Ease Score (FRES), and Flesch-Kincaid Grade Level (FKGL) were calculated. These indices estimate reading difficulty or grade level and provide quantitative measures of text readability (Table 2).
Evaluation of educational suitability and information quality
Educational suitability was assessed using the Patient Education Materials Assessment Tool (PEMAT), which comprises two domains: understandability and actionability28. The instrument includes 24 items, with 16 evaluating understandability and eight evaluating actionability. Each item was scored dichotomously (0 = criterion not met; 1 = criterion met), yielding a total score ranging from 0 to 24, with higher scores indicating greater suitability for patient education. Overall information quality was evaluated using the Global Quality Score (GQS)27, a widely used five-point Likert scale for assessing the quality of medical information (Table 3).
On November 25, 2025, two clinical experts with more than 3 years of experience in TN management evaluated all AI-generated responses using both assessment instruments. Before scoring, all AI-generated responses were anonymized with coded response identifiers, and reviewers were blinded to the LLM platform during PEMAT and GQS assessments. Disagreements were resolved by a third senior expert, who adjudicated the final scores. Inter-rater reliability was assessed using Cohen's kappa coefficient, with values >0.75 indicating excellent agreement, 0.40–0.75 indicating acceptable agreement, and <0.40 indicating poor agreement. Cohen's kappa values for both the PEMAT and GQS exceeded 0.75, demonstrating excellent inter-rater reliability.
Statistical analysis
All statistical analyses were performed using R software (version 4.4.3). Data normality was assessed using the Shapiro-Wilk test and Q-Q plots. Normally distributed variables were summarized as mean ± standard deviation and compared using one-way analysis of variance (ANOVA), followed by Tukey's post hoc test when appropriate. Non-normally distributed variables were summarized as medians and interquartile ranges and compared using the Kruskal-Wallis test, followed by Dunn's post hoc test with a Bonferroni correction for pairwise comparisons. Correlations among readability indices, PEMAT scores, and GQS values were evaluated using Spearman's rank correlation coefficient, with the Benjamini-Hochberg correction applied for multiple-testing. A two-sided adjusted P value < 0.05 was considered statistically significant.
Access restricted. Please log in or start a trial to view this content.
This study systematically evaluated the effects of large language models (LLMs) and different categories of health education content on the readability and quality of generated texts. First, we compared the performance of five LLMs—Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and GPT-5—with respect to patient education suitability, overall information quality, and multiple readability metrics, including the PEMAT score, GQS, and established readability indices (ARI, FRES, GFOG, FKGL, CL, SMOG, and LW). Second, we exam...
Access restricted. Please log in or start a trial to view this content.
This cross-sectional benchmarking study systematically compared the readability, educational suitability, and overall information quality of trigeminal neuralgia (TN) patient education materials generated by five publicly available large language models (LLMs). Four principal findings emerged. First, readability differed substantially across the models, with GPT-5 generating the most linguistically complex responses and Wenxin Yiyan producing the most readable content. Second, readability and expert-rated information qua...
Access restricted. Please log in or start a trial to view this content.
The authors declare that they have no competing interests.
This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. The authors thank the clinical colleagues who provided feedback on the trigeminal neuralgia question set and assisted in refining the evaluation framework. We also thank all contributors who supported manuscript preparation and revision.
Access restricted. Please log in or start a trial to view this content.
| Name | Company | Catalog Number | Comments |
|---|---|---|---|
| DeepSeek chat platform | DeepSeek | https://chat.deepseek.com/ | Publicly accessible large language model interface used to generate responses |
| Doubao chat platform | Doubao | https://www.doubao.com/chat/ | Publicly accessible large language model interface used to generate responses |
| GPT 5 | OpenAI | https://chat.openai.com/ | Publicly accessible large language model interface used to generate responses |
| R software, version 4.4.3 | R Foundation for Statistical Computing | Not applicable | Statistical analysis and visualization |
| Readability Formulas online readability calculator | Readability Formulas | Not applicable | Online tool used to compute readability indices |
| Tongyi Qianwen chat platform | Tongyi Qianwen | https://www.tongyi.com/ | Publicly accessible large language model interface used to generate responses |
| Wenxin Yiyan chat platform | Wenxin Yiyan | https://yiyan.baidu.com/ | Publicly accessible large language model interface used to generate responses |
Access restricted. Please log in or start a trial to view this content.