Eligibility criteria
Studies were eligible if they met all of the following criteria: (1) included women with any gynecologic malignancy; (2) evaluated a defined lifestyle intervention or a nursing‑led intervention; (3) reported quantitative quality‑of‑life data; and (4) were either comparative (intervention vs. control) or observational trajectory studies that provided relevant descriptive QoL data (5) were published in English. . For the quantitative meta‑analysis of intervention versus control, a concurrent comparator group was required12. Two studies without a comparator were retained for narrative synthesis only13. No geographic restrictions were applied.
Information sources
The full study selection process is shown in Figure 1 as a PRISMA flow diagram. The Cochrane Library, Embase, Google Scholar, OVID, and PubMed were searched from database inception through June 2026. Retrieved records were imported into EndNote 20, and duplicate records were removed. Two reviewers independently screened titles and abstracts, followed by full-text assessment of potentially eligible studies. Disagreements were resolved by discussion. Only English-language publications were considered for inclusion, consistent with the prespecified eligibility criteria.
Search strategy
The search strategy was developed using the PICOS framework: Population—women with gynecologic oncology; Intervention—lifestyle or nursing interventions; Comparison—usual care or alternative interventions; Outcome—quality of life; Study design—no restrictions14. A comprehensive set of keywords and controlled vocabulary terms was used across all databases; the detailed search strings for each database are provided in Table 113.
Selection process
Following full‑text retrieval, two reviewers independently verified that each report met all eligibility criteria. They additionally cross‑checked author lists, trial names (e.g., SUCCEED, WALC, BENITA, TOPCAT‑G), recruitment periods, and study centers to identify overlapping patient cohorts. Potentially overlapping patient cohorts were identified and documented. When overlapping reports contributed to the primary analysis, their influence was evaluated in a prespecified sensitivity analysis that excluded the overlapping dataset.
Data extraction and outcome measures
Two reviewers independently extracted data from all 15 included reports using a standardized, piloted data‑extraction form. Disagreements were resolved by consensus with the corresponding author. For each study, the following information was recorded: first author, publication year, country, study design, cancer type, intervention description, comparator description, sample size per group, participant age and disease stage, quality‑of‑life (QoL) instrument, time points assessed, and all numerical QoL data (means, measures of dispersion, and sample sizes per arm)15.
Quantitative pooling (13 comparative studies)
For the meta‑analysis of intervention versus control, post-intervention final scores (means and standard deviations) were extracted at the prespecified time points (3 and 6 months) for both groups. Where a study reported both final and change scores, only final scores were extracted16. The methodological approach to meta-analysis and assessment of study quality followed established Cochrane guidance17. The choice of meta-analytic model considered anticipated clinical and methodological heterogeneity across studies18, and statistical heterogeneity was quantified using the I2 statistic as described by Higgins et al.19. The 13 comparative studies included in the quantitative synthesis are summarized in Table 220,21,22,23,24,25,26,27,28,29,30,31,32.
Quality‑of‑life outcome measures, alignment, and effect metric
The specific QoL instrument used in each study was recorded. The identified scales included the Functional Assessment of Cancer Therapy – General (FACT‑G), FACT‑Ovarian (FACT‑O), FACT‑Endometrial (FACT‑En), the European Organization for Research and Treatment of Cancer Quality of Life Questionnaire Core 30 (EORTC QLQ‑C30), the Short Form‑36 (SF‑36), and condition‑specific scales. For all instruments, higher scores indicated better quality of life, and no reverse‑scored instruments were included. All extracted QoL data were therefore aligned to the same metric, with a positive mean difference (MD) consistently favoring the intervention group.
For the primary meta‑analysis, post-intervention final scores were extracted rather than change‑from‑baseline scores. This decision was based on: (i) not all studies reported change scores or their standard deviations; (ii) final scores reflect the absolute post‑treatment QoL status, which is clinically meaningful; and (iii) using final scores maximized the number of studies that could be pooled. When a study reported multiple QoL subscales, the global QoL score or total FACT-G score was prioritized as the primary outcome; if unavailable, the physical-function subscale was used, with this deviation explicitly footnoted in Table 2.
The raw MD was used as the effect measure rather than the standardized mean difference (SMD) because all included instruments are validated to measure the same underlying construct—health-related quality of life—on conceptually similar scales (total scores range 0–100 or 0–148). The MD is directly interpretable in the original units of these scales and is more clinically meaningful than an SMD expressed in standard‑deviation units. Adjusted estimates (e.g., ANCOVA‑adjusted means) were not used because the adjustment covariates varied widely across studies, compromising comparability.
Standard deviation conversions
When studies reported standard errors (SEs) or 95% confidence intervals instead of standard deviations, they were converted to SDs using standard formulas: SD = SE × √n for SEs, and SD = √n × (upper CI – lower CI) / (2 × 1.96) for 95% CIs. For studies presenting medians with interquartile ranges (IQR), means, or SDs were not imputed; such data were extracted narratively and were not included in the pooled MD calculations. No such imputation was required for the 13 studies that contributed to the quantitative synthesis.
Missing data handling
Intention-to-treat (ITT) estimates were preferentially extracted when explicitly reported, as they provide the most conservative and unbiased estimate of the treatment effect. Where ITT data were unavailable, per-protocol or complete-case data were extracted as reported by the original authors. Missing means, SDs, or outcome values were not imputed; missingness was recorded for each study and accounted for in the sensitivity analysis. Studies with unreported SDs for key time points were noted, and the corresponding authors were contacted by email to request the missing statistics. If no response was received within two weeks, the most conservative available data were used (e.g., pooled baseline SD if post‑intervention SD was missing) or the specific time point was excluded from pooling.
Multiple intervention arms and shared controls
No included study presented more than one active lifestyle/nursing intervention arm compared with a single shared control group, which would have required splitting the control sample for analysis. In the one factorial‑designed study (McCloy et al.30), data were extracted from the combined exercise-and-mindfulness arm and the usual‑care control arm to simplify pooling and avoid double‑counting. For the overlapping SUCCEED trial cohorts (Von Gruenigen et al.21 and McCarroll et al.23), both reports were retained in the extraction table. For the primary meta‑analysis, both datasets were included (as they reported QoL at the same time point). In a prespecified sensitivity analysis, the latter report was excluded to assess the impact of cohort duplication on the pooled estimate; results are reported in the text.
Non‑comparative observational studies
For two reports (Budisan et al.33 and Betea et al.34), which lacked a concurrent control group, intervention-versus-control comparative data were not extracted (i.e., no MD could be computed). Instead, descriptive QoL trajectory data were extracted (within‑group changes over time) and presented narratively in the text and in Table 2. They were not included in any forest plot or pooled quantitative synthesis.
Risk of bias assessment
Potential publication bias and other small‑study effects were assessed using both visual inspection of funnel plots and Egger's linear regression test (which regresses the standardized intervention effect on its precision). Because Egger's test has limited statistical power when the number of studies is small (generally <10), funnel plots and Egger's tests were prespecified and would be performed only for meta‑analyses comprising 10 or more studies. For outcomes with fewer than 10 contributing studies, funnel plots and Egger p-values are not reported, as such tests are underpowered and may yield misleading results17.
Two reviewers independently assessed the methodological quality of the 15 included reports, and disagreements were resolved through discussion with the corresponding authors. Because the included studies comprised both randomized controlled trials (RCTs; n = 13) and non‑randomized observational cohort studies (n = 2; Budisan et al.33; Betea et al.34), design-specific tools were applied.
For RCTs (13 studies), the Cochrane RoB 2.0 tool was used (revised version), evaluating each study across five domains17: (1) Randomization process -- sequence generation and allocation concealment; (2) Deviations from intended interventions -- blinding of participants and personnel, and whether the analysis was intention‑to‑treat; (3) Missing outcome data -- completeness of follow‑up and handling of dropouts; (4) Measurement of the outcome -- blinding of outcome assessors (particularly relevant for self‑reported QoL, where blinding is less feasible but still considered); (5) Selection of the reported result -- risk of selective outcome reporting (checked against trial registries or published protocols where available).
Each domain was rated as low risk, some concerns, or high risk. An overall risk‑of‑bias judgment was then assigned for each study: low (all domains low), some concerns (one or more domains with some concerns), or high (any domain rated high risk or multiple domains with some concerns). For self‑reported outcomes such as quality of life, blinding participants is often infeasible; a high-risk judgment was therefore not automatically assigned for lack of participant blinding. Instead, lack of blinding was explicitly noted as a potential source of performance bias in the domain 2 assessment.
For non-randomized observational cohort studies (2 studies), Budisan et al.33 and Betea et al.34 did not include a concurrent comparator group and were not pooled in the quantitative meta‑analysis. For these, the Newcastle-Ottawa Scale (NOS) was used, adapted for cohort studies, to assess selection, comparability (not applicable due to the lack of a comparator), and outcome ascertainment. However, because they lacked a control arm, they were judged to be at high risk of confounding and selection bias in any causal inference. They were retained only for descriptive purposes; their risk-of-bias profiles do not affect the pooled effect estimates, but they are documented in Table 2.
Summary of risk of bias
A risk-of-bias summary figure (traffic-light plot) was created, displaying domain‑level judgments for each RCT. In brief, the most common concerns across RCTs were: (i) lack of blinding of participants and personnel to lifestyle interventions (domain 2 – some concerns/high risk in 9 of 13 studies), and (ii) incomplete outcome data due to attrition (domain 3 – some concerns in 4 studies). No study was rated as low risk across all domains; the majority (8/13) had some concerns, and 5/13 were rated high risk overall, primarily due to lack of blinding and small sample sizes leading to imprecision.
Certainty of evidence (GRADE)
The overall certainty of the body of evidence was assessed for each pooled outcome using the Grading of Recommendations Assessment, Development and Evaluation (GRADE) framework, following the GRADE Handbook and using GRADEpro GDT software. Certainty was rated as high, moderate, low, or very low across five domains: risk of bias, inconsistency, indirectness, imprecision, and publication bias. Outcomes assessed include (1) QoL at 6 months – general Gynecologic Oncology (pooled MD from 9 comparative studies); (2) QoL at 6 months – ovarian cancer (pooled MD from 2 comparative studies); (3) QoL at 6 months – endometrial cancer (pooled MD from 4 comparative studies, including overlapping SUCCEED cohorts); (4) QoL at 3 months – general Gynecologic Oncology (pooled MD from 6 comparative studies).
Effect measure and meta‑analysis model
Random-effects inverse-variance models were used for the general 6-month and 3-month analyses because substantial clinical and methodological heterogeneity was anticipated across studies, including differences in cancer types, intervention components, intensity, duration, and control conditions18. Fixed-effect inverse-variance models were used for the ovarian and endometrial cancer subgroup analyses, for which no statistical heterogeneity was observed (I2 = 0%). Heterogeneity was assessed using the I2 statistic19.
Heterogeneity was quantified using the I2 statistic, with thresholds of 25%, 50%, and 75% representing low, moderate, and high heterogeneity, respectively. In addition to I2, τ2 was reported for random-effects analyses to quantify the absolute magnitude of between-study variance. All analyses were performed using the meta and metafor packages in R (version 4.3.1).
Justification for pooling despite substantial heterogeneity and handling of heterogeneity
For the primary 6-month analysis, I2 = 71% indicated substantial heterogeneity. Pooling was considered appropriate under the random-effects model for the following reasons: (i) the pooled estimate represents the average effect across a range of interventions and populations, which is a legitimate summary question; (ii) the random-effects model explicitly incorporates between‑study variance (τ2) into the standard error, reducing the risk of falsely narrow confidence intervals; (iii) prespecified subgroup analyses were conducted by cancer type (ovarian, endometrial) to explain clinical sources of heterogeneity; (iv) influence and sensitivity analyses were performed to assess the robustness of the pooled estimate.
The decision to pool studies despite substantial statistical heterogeneity was guided by the following principles: I2 interpretation — I2 values of 25%, 50%, and 75% were prespecified to represent low, moderate, and high heterogeneity, respectively. An I2 exceeding 75% indicates considerable heterogeneity, which may make pooling inappropriate if the heterogeneity cannot be meaningfully explained by clinical or methodological factors. For the primary 6‑month general analysis (I2 = 71%), subgroup analyses by cancer type (ovarian, endometrial) were prespecified to explore clinical sources of heterogeneity. Subgrouping substantially reduced I2 to 0% in both ovarian and endometrial subgroups, suggesting that a major portion of the overall heterogeneity is attributable to differences in cancer type and associated intervention targets. However, because these subgroups contain very few studies, it cannot be reliably concluded that heterogeneity is fully explained.
Proceeding with pooling was considered justified for the general 6‑month analysis because: (i) it provides a summary average effect across a broad range of interventions, which is a legitimate meta‑analytic question; (ii) τ2 is explicitly reported to quantify between‑study variance; (iii) influence analyses were performed to check for outlier‑driven results; and (iv) the pooled estimate is not treated as definitive or clinically actionable — it is interpreted as hypothesis‑generating and is heavily caveated in the discussion.
In alignment with the Cochrane Handbook, when I2 exceeds 75%, and subgroup analyses do not convincingly explain the heterogeneity, pooled estimates should be interpreted with extreme caution. Given the small number of studies in the subgroups, it cannot be confidently asserted that heterogeneity is fully explained. Therefore, the primary pooled estimate is presented primarily for descriptive and exploratory purposes and does not form the basis for clinical recommendations. All heterogeneity statistics (I2, τ2) and sensitivity analyses are reported transparently, allowing readers to assess the stability of the estimates themselves.
Subgroup, sensitivity, and influence analyses
To explore sources of heterogeneity, the following prespecified analyses were performed, without conducting meta‑regression: 1) For subgroup analyses, the general 6-month outcome was stratified by cancer type (ovarian, endometrial) to examine whether the effect differed by tumor site; 2) Sensitivity analysis (overlapping cohorts): To assess the impact of the overlapping SUCCEED trial cohorts (Von Gruenigen et al.21 and McCarroll et al.23), the endometrial subgroup analysis was repeated excluding McCarroll et al.23. 3) Leave‑one‑out influence analysis: Each study was sequentially removed, and the pooled MD and I2 were recalculated for the general 6‑month outcome to identify whether any single study disproportionately drove the results or the heterogeneity; 4) Fixed‑effect model comparison: For completeness, pooled estimates were also computed under a fixed‑effect model (inverse‑variance weighting) to compare with the random‑effects results; any substantial discrepancy would indicate that between‑study heterogeneity meaningfully influences the estimate.
Meta-regression was not performed to explore continuous covariates (e.g., age, baseline QoL, intervention duration) because the number of studies per subgroup (≤9 for the general 6‑month outcome) is insufficient to yield reliable regression coefficients. With <10 studies per covariate, meta‑regression is highly underpowered and prone to type I errors, and the available covariate data were inconsistently reported across studies.
Publication bias
None of the pooled outcomes included 10 or more studies. Therefore, funnel plots and Egger's regression tests were not performed because these assessments would be underpowered and potentially misleading given the number of studies available.