$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
The NHANES protocol was approved by the National Center for Health Statistics (NCHS) Research Ethics Review Board, and written informed consent was obtained from all participants. This work was a secondary analysis of de-identified public-use data; therefore, no additional institutional ethical approval was required. All authors read and approved the final manuscript.
1. Study design and data source
This study was conducted as a secondary analysis of NHANES, a series of cross-sectional, nationally representative surveys administered by the U.S. Centers for Disease Control and Prevention and overseen by the NCHS Research Ethics Review Board. Public-use NHANES datasets were fully de-identified and were accessed for secondary analysis. Data from the 1999–2000, 2001–2002, 2003–2004, and 2005–2006 cycles were used.
Public NHANES component files were downloaded for each cycle, including (i) demographic files containing the participant identifier (SEQN) and survey design variables, (ii) reproductive health questionnaire files containing endometriosis self-report, and (iii) laboratory files required for the composite index computation (C-reactive protein, triglycerides, and fasting plasma glucose). Examination/anthropometry files (e.g., body mass index) and laboratory measures required for comparator indices (e.g., neutrophils, lymphocytes, platelets) were additionally obtained when those indices were analyzed. Within each 2-year cycle, component files were merged using SEQN, and the merged dataset was checked to ensure one record per SEQN. Cycle-level datasets were then appended to construct the combined 1999–2006 analytic file.
The analytic sample was restricted to women aged 20–54 years. Participants were excluded if endometriosis status was missing, if any composite-index components (C-reactive protein, triglycerides, or fasting plasma glucose) were missing, or if essential covariates required for the fully adjusted model were missing under a complete-case strategy. Complex survey design variables (strata and primary sampling units) were retained, along with the fasting laboratory subsample weights required for analyses incorporating fasting measures. When multiple NHANES cycles were combined, multi-cycle weights were created according to NHANES analytic guidance by dividing the 2-year subsample weight by the number of combined cycles, and the resulting weight, strata, and PSU variables were applied in all analyses. The steps for participant inclusion and exclusion were documented in a flow diagram (Figure 1).
2. Definition of endometriosis
Endometriosis status was defined using the reproductive health questionnaire item: “Have you ever been told by a doctor or other health professional that you have endometriosis?” Participants who responded “Yes” were classified as endometriosis cases, and those who responded “No” were classified as controls. Because this definition was based on self-report rather than laparoscopic or histologic confirmation, potential misclassification was addressed as a study limitation.
3. Definition of the Composite Index (CTI)
The C-reactive protein–triglyceride–glucose composite index was operationalized to jointly reflect systemic inflammation and metabolic disturbance. Laboratory measurements of C-reactive protein (mg/L), triglycerides (mg/dL), and fasting plasma glucose (mg/dL) were extracted from NHANES laboratory files. The triglyceride–glucose index was calculated as the natural logarithm of [triglycerides × fasting plasma glucose/2]. CTI was calculated using the following formula: CTI = 0.412 × ln(CRP) + TyG. Higher CTI values indicate a higher combined burden of low-grade inflammation and insulin resistance8.
If any C-reactive protein values required handling prior to log transformation (e.g., values at or below the detection limit), a single prespecified rule was applied consistently across all cycles and was documented to support replicability (for example, replacing non-positive values with the smallest positive measurable value observed prior to log transformation). Quartile cut points were determined from the weighted distribution in the full analytic sample and were applied consistently across categorical analyses, with Quartile 1 used as the reference category.
4. Covariates
Covariates were prespecified to mitigate confounding based on epidemiologic reasoning and prior literature. Demographic variables included age, race/ethnicity, education level, and marital status. Lifestyle variables included smoking history (≥100 cigarettes in lifetime vs. <100) and alcohol consumption (≥12 drinks/year vs. <12). Comorbidity history included self-reported hypertension, diabetes, stroke, coronary heart disease, and cancer. Anthropometric and laboratory variables included body mass index, hemoglobin, neutrophil count, lymphocyte count, and platelet count; these measures also supported computation of comparator inflammatory indices where applicable (e.g., neutrophil-to-lymphocyte ratio, platelet-to-lymphocyte ratio, systemic immune-inflammation index, systemic inflammation response index). Reproductive variables (e.g., gravidity and parity) were included when available in the selected cycles and were coded according to NHANES documentation. Categorical covariates were converted to indicator variables prior to model entry.
Because C-reactive protein was a component of the composite index, it was not entered as an independent covariate in multivariable regression models to avoid overadjustment and collinearity. Instead, C-reactive protein and the triglyceride–glucose index were evaluated as comparator markers in the discrimination analyses.
5. Statistical analysis
All analyses accounted for the NHANES complex survey design to generate nationally representative estimates. The survey design was specified by linking the multi-cycle subsample weight, strata, and PSU variables to the analytic dataset. Continuous variables were summarized as weighted means with standard deviations, and categorical variables were summarized as weighted counts and percentages. Baseline characteristics were compared between cases and controls using survey-weighted procedures appropriate for NHANES, and baseline characteristics were summarized in Table 1.
Associations between the composite index and endometriosis were evaluated using survey-weighted logistic regression. Three sequential models were fitted to demonstrate adjustment: an unadjusted model, a model adjusted for age and race/ethnicity, and a fully adjusted model including demographic factors, lifestyle variables, comorbidity history, anthropometric/laboratory covariates, and reproductive history variables. The composite index was analyzed both continuously (per 1-unit increase) and categorically (quartiles, with Quartile 1 as the reference), and regression estimates were summarized in Table 2. Linear trend across quartiles was tested by assigning each quartile its weighted median value and modeling that term continuously.
Non-linear dose–response relationships were assessed using survey-weighted restricted cubic splines with prespecified knot placement, and spline curves were plotted in Figure 2. Threshold effects were evaluated using survey-weighted segmented (piecewise) logistic regression by comparing model fit between segmented and single-slope specifications, and the estimated inflection point and slope parameters on each side of the inflection point were reported in Table 3.
Subgroup analyses were conducted to explore effect modification by prespecified factors (e.g., age group, race/ethnicity, education level, marital status, and selected lifestyle factors). Interaction was tested by including cross-product terms between the continuous composite index and subgroup indicators within the survey-weighted framework, and subgroup associations were summarized in Figure 3.
Discriminatory performance was evaluated using receiver operating characteristic analyses based on model-predicted probabilities derived from survey-weighted logistic models. Area under the curve estimates were obtained for the composite index, commonly used inflammatory indices, and component markers, and AUC summaries were provided in Supplementary Table 1; an expanded ROC comparison was provided in Supplementary Figure 1.
Missing data were handled using complete-case analysis after excluding participants with missing endometriosis status, missing composite-index components, or missing essential covariates required for the fully adjusted model. When a robustness assessment was performed, multiple imputation was applied for covariates with missingness under a prespecified imputation model, and imputed estimates were compared with complete-case estimates.
Analyses were performed using R (version 4.4.1) and additional statistical software as listed in the Table of Materials. Key packages used for survey inference, spline modeling, segmented regression, and ROC estimation were recorded, and session information (operating system and R session details) was retained to support replication.
6. Procedure end point and outputs
The analytic workflow was considered complete once the harmonized multi-cycle dataset was constructed with the prespecified inclusion/exclusion criteria (Figure 1), the composite index and covariates were generated according to documented rules, and the prespecified survey-weighted regression, non-linearity/threshold assessment, subgroup, and discrimination analyses were executed under the same survey design specification. The primary outputs of this workflow were organized as a baseline summary (Table 1), regression estimates across sequential adjustment models (Table 2), threshold model parameters (Table 3), spline visualization (Figure 2), subgroup summary visualization (Figure 3), and discrimination summaries (Supplementary Table 1 and Supplementary Figure 1).
7. Internal independent validation
Internal independent validation was performed by splitting the combined dataset into a derivation cohort and a non-overlapping validation cohort based on NHANES cycles. Participants from the 1999–2000 and 2001–2002 cycles were assigned to the derivation cohort, and participants from the 2003–2004 and 2005–2006 cycles were assigned to the validation cohort. The same inclusion/exclusion criteria, composite index computation, covariate coding rules, and survey-weighting strategy were applied independently within each cohort.
Within the derivation cohort, survey-weighted logistic regression models were fitted using the fully adjusted specification. The non-linearity and threshold assessment procedures used in the main analysis were applied in the derivation cohort, and discrimination was evaluated using ROC/AUC methods based on model-predicted probabilities. The same modeling strategy was then repeated in the validation cohort without modifying variable definitions, coding rules, or weighting specifications. Derivation-versus-validation estimates were summarized as a cohort-to-cohort comparison of association estimates (Figure 4) and as derivation-versus-validation ROC curves (Figure 5), with corresponding numerical summaries provided in Table 4.