$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
This study was conducted in strict accordance with the ethical principles outlined in the Declaration of Helsinki. The study protocol was reviewed and formally approved by the Institutional Review Board (IRB) of Chongqing Medical University (ethical approval no.: 2026127). Prior to enrollment, all participants were provided with a comprehensive explanation of the study’s objectives, procedures, and potential risks, and subsequently provided written informed consent. To safeguard patient privacy and ensure data confidentiality, all collected information was fully anonymized and de-identified before being accessed for statistical analysis. The complete list of survey instruments, expert consultation forms, and analytical software utilized in this study is detailed in the Table of Materials.
Study design and participant recruitment
A multi-stage methodological study, comprising a cross-sectional baseline survey and subsequent model-driven predictive simulations, was conducted across five primary healthcare institutions. Participants were recruited utilizing a stratified convenience sampling strategy during their routine follow-up visits. Eligibility criteria included: (1) adults aged 18 years or older; (2) medically diagnosed with at least one chronic condition (e.g., hypertension, type 2 diabetes) managed at the participating institution for a minimum of six months; and (3) possessing the cognitive ability to independently complete the evaluation. The surveys were administered face-to-face by trained clinical staff. To address missing data, Little's MCAR test was initially applied. Given that the missing data rate was low (<5%) and missing completely at random, multiple imputation techniques were employed to handle incomplete responses, ensuring the integrity of the dataset. To avoid overfitting and clarify the sample basis for each analytical stage, the fully imputed dataset (N = 433) was strictly partitioned. Specifically, 50% of the sample (n = 217) was utilized for exploratory factor analysis (EFA) to establish the factor structure. The remaining 50% (n = 216) was reserved for confirmatory factor analysis (CFA), structural equation modeling (SEM), and baseline establishment for institution-level quality profiling. The comparative model validation and simulated optimization scenarios were subsequently derived strictly from this second hold-out sample (n = 216) to prevent data leakage.
Conceptual framework and indicator development
The construction of the evaluation index system began by mapping the five core SERVQUAL dimensions—tangibility, reliability, responsiveness, assurance, and empathy—to the specific workflow of primary chronic disease management. This process involved integrating concrete service touchpoints, such as medical record creation, pharmaceutical distribution, and longitudinal follow-up protocols. To ensure relevance for the target demographic, the primary care information system was analyzed to assess its accessibility for elderly patients, followed by iterative refinements to the items. Specific reliability metrics focused on the accuracy of follow-ups and staff commitment, while responsiveness was operationalized as the timeliness of responses to patient inquiries (see Table 1 for operational definitions). The overall framework, including the interrelationships between dimensions, is conceptualized in Figures 1 and 2.
Delphi expert consultation and content validity
A two-round Delphi method was employed to refine the initial item pool. Experts were purposively recruited based on a multi-dimensional set of eligibility criteria to ensure professional authority and methodological rigor. Candidates were invited to the panel if they met the following requirements: (1) holding a senior professional title (e.g., Associate Professor, Chief Physician, or Nursing Director); (2) possessing a minimum of ten years of clinical, managerial, or research experience specifically focused on primary healthcare or chronic disease management; and (3) demonstrating a recognized track record in health service evaluation or public health policy development. The final panel of fifteen experts represented a cross-disciplinary cohort, including primary care physicians (n = 5), nursing managers (n = 4), public health researchers (n = 4), and health policy administrators (n = 2). This diverse composition and the rigorous selection criteria directly ensured the reproducibility and credibility of the subsequent content validity and AHP weighting procedures. Experts evaluated each item's relevance and clarity using a 4-point Likert scale (1 = not relevant, 4 = highly relevant). Consensus was defined a priori as ≥80% agreement among experts rating an item as 3 or 4. Items with an Item-Level Content Validity Index (I-CVI) ≥ 0.78 were retained (Table 2). Disagreements between rounds were resolved through iterative anonymous feedback; experts were provided with aggregated scores from the previous round and permitted to adjust their ratings. Items failing to reach consensus after the second round were either substantially revised based on qualitative expert feedback or eliminated, culminating in the final scale-level content validity index (S-CVI)12,13.
Analytic hierarchy process (AHP) and scoring mechanism
To determine the relative weights of the evaluation indicators, the AHP method was utilized. A purposive sub-sample of nine senior experts from the Delphi panel completed a 9-point fundamental scale to construct pairwise comparison matrices for all dimensions and sub-dimensions. The geometric mean method was used to aggregate the experts' judgments. A Consistency ratio (CR) was calculated for each matrix; only matrices with a CR of were considered acceptable, ensuring logical consistency in expert judgments. Final global weights (Wi) were derived by multiplying the local weights of sub-dimensions by the weights of their parent dimensions. The final service quality score (SQ) for each patient was calculated using the formula:
SQ = ∑(Wi x Pi) (1)
where Pi represents the patient's perceived performance score for the item i. Notably, this study adopted a perception-only (P) scoring method—consistent with the SERVPERF paradigm—rather than the traditional gap-based (P-E) approach. This methodological choice, measured on a 7-point Likert scale (ranging from 1 = "Strongly Disagree" to 7 = "Strongly Agree"), was implemented to enhance the model's predictive validity and minimize the measurement instability often associated with quantifying patient expectations in longitudinal chronic care settings.
Psychometric validation and structural equation modeling
The randomly split dataset was subjected to psychometric evaluation. For the first half of the data, EFA was conducted using principal axis factoring with promax oblique rotation, as the SERVQUAL dimensions were hypothesized to be correlated. Factors were retained based on eigenvalues >1 and visual inspection of the scree plot. Items with factor loadings <0.40 or significant cross-loadings were excluded. The second half of the data was used for CFA to confirm the factor structure. Maximum likelihood (ML) estimation was employed for the CFA and subsequent SEM analyses. Model fit was evaluated using stringent thresholds established in the psychometric literature, which indicate an acceptable balance between model parsimony and data fit: Comparative fit index (CFI) > 0.90, Tucker-Lewis index (TLI) > 0.90, root mean square error of approximation (RMSEA) < 0.08, and standardized root mean square residual (SRMR) > 0.08.
Predictive model comparison plan
To assess the comparative effectiveness of different evaluative frameworks, a predictive modeling analysis was conducted. To establish the methodological robustness of the proposed framework, the basic SERVQUAL model was benchmarked against three distinct analytical approaches: the proposed AHP-SERVQUAL, the KANO model, and the TOPSIS-RSR (Technique for Order Preference by Similarity to Ideal Solution combined with Rank-Sum Ratio) method. The KANO model was selected as a comparator because it categorizes service attributes into non-linear satisfaction drivers (e.g., basic needs versus attractive qualities), offering a theoretical contrast to the linear assumptions of SERVQUAL. Conversely, TOPSIS-RSR is a rigorous multi-criteria decision-making (MCDM) tool widely utilized for comprehensive institutional ranking, thereby serving as a robust methodological control for our AHP-based scoring system. To execute the comparison, each of the four models was independently applied to the hold-out validation dataset (n = 216), incorporating identical context-specific parameters (such as digital follow-up and self-management support). The evaluative process subsequently compared how effectively each model's generated quality indices correlated with external performance metrics (complaint frequency and issue resolution rates) and their predictive power (R2) for patient compliance. The predictive accuracy of each model was quantified using R2 values generated through linear regression and SEM. To ensure model parsimony and avoid over-fitting, information criteria, specifically the Akaike information criterion (AIC) and Bayesian information criterion (BIC), were integrated into the selection process. Finally, cross-validation was executed to evaluate the robustness of the selected models across the five participating institutions. The comparative metrics for problem resolution, complaint frequency, and overall predictive power (R2) are shown in Figures 3 and 4.