Eligibility criteria
Studies were included if they met all of the following criteria, organized according to the PICOS framework (Population, Intervention, Comparator, Outcomes, Study design)17. All platforms and materials used in this study are listed in the Table of Materials. All database searches were performed in November 2025. No additional proprietary software (e.g., SAS, SPSS, R) was used beyond RevMan 5.3 and EndNote 20. The GRADE assessment was performed manually using the GRADEpro GDT online tool; the version is recorded as of the access date.
Population (P)
Participants had clinically significant depressive symptoms, defined as either a formal diagnosis of major depressive disorder (MDD) or dysthymia according to DSM criteria (any version) or ICD criteria, confirmed by structured clinical interview (e.g., SCID, MINI, CIDI); or a score above the validated cutoff on a standardized depression screening instrument (e.g., Beck Depression Inventory ≥ 14, Hamilton Depression Rating Scale ≥ 14, CES-D ≥ 16, PHQ-9 ≥ 10, Geriatric Depression Scale ≥ 11). Participants could be of any age (children, adolescents, adults, older adults)18. Studies that enrolled participants with comorbid medical or psychiatric conditions (e.g., cardiovascular disease, diabetes, anxiety disorders, substance use disorders) were included only if the primary analysis reported results separately for the depression subgroup or if ≥80% of participants met depression criteria. Studies where depression was a secondary outcome in a non-depressed population (e.g., heart failure patients without depression screening) were excluded. There was no restriction on baseline depression severity (mild, moderate, severe), but severity was extracted for subgroup analyses.
Intervention (I)
The intervention consisted of structured, planned physical exercise, defined as any form of aerobic, resistance, mixed (aerobic + resistance), or alternative exercise (yoga, tai chi, dance) delivered for at least 2 weeks. Exercise interventions could be supervised (by a trainer, therapist, or instructor) or unsupervised (home-based, self-directed). Studies where exercise was combined with another active intervention (e.g., exercise + pharmacotherapy, exercise + mindfulness) were included only if the effect of exercise could be isolated (e.g., exercise + pharmacotherapy vs. pharmacotherapy alone). Studies where the intervention consisted only of physical activity advice (without structured exercise prescription) or lifestyle counseling were excluded.
Comparator (C)
The control group received no structured exercise intervention. Acceptable control conditions were categorized as passive controls (no active therapeutic intervention) e.g. waitlist (delayed treatment), usual care (standard medical management without structured exercise), or no intervention; and active controls (non‑exercise therapeutic interventions) e.g. attention control (e.g., non‑exercise health education, relaxation sessions, or light stretching not meeting exercise intensity criteria), pharmacotherapy (antidepressant medication alone), or psychotherapy (psychological therapy without exercise). Studies in which the control group received any form of structured exercise (including minimal exercise or stretching that met the exercise intensity criteria) were excluded. Note on discrepancy, although the innovation section stated that studies with active psychological interventions or stretching/relaxation controls were excluded, subsequently, the eligibility criteria to include active comparators (pharmacotherapy, psychotherapy, attention control) to reflect the real‑world clinical question of how exercise compares to other established treatments were broadened. This discrepancy is addressed in the discussion. All analyses stratify by control type (passive vs. active) to examine the influence of control type on effect sizes.
Outcomes (O)
The study reported at least one depression-specific outcome measure: depression remission (dichotomous, defined by study authors as no longer meeting diagnostic criteria or scoring below the cutoff on a validated scale); OR (b) change in continuous depression severity score measured on a validated scale (e.g., HAM-D, BDI, CES-D, PHQ-9, GDS, MADRS). Studies reporting only non-specific mood measures (e.g., “well-being,” “emotional state” without a validated depression scale) were excluded.
Study design (S)
Only randomized controlled trials (RCTs) were included. Prospective cohort studies, case-control studies, retrospective studies, case series, and case reports were excluded because they are at high risk of confounding and selection bias when estimating treatment effects. Studies could be parallel-group or crossover design (only the first phase extracted for crossover trials). Studies were required to have at least 10 participants per arm at baseline to minimize small-study effects. No language restrictions were applied, but non-English studies were included only if a full translation was available.
Exclusion criteria
Studies where depression was not the primary outcome or where the population was not selected for depression (e.g., general medical populations with depression as a secondary measure) were excluded. Studies in which the exercise intervention was delivered as part of a multi-component lifestyle intervention, without an isolated exercise effect, were excluded. Studies where the control group received any form of structured exercise were excluded. Studies with insufficient data to calculate effect sizes (means, standard deviations, or change scores not reported and not obtainable from authors) were excluded. Duplicate publications of the same study cohort (only the most complete or most recent report was included) were excluded. Conference abstracts, dissertations, and unpublished manuscripts without full-text availability were excluded.
Information sources
The full study is shown in Figure 1 as a PRISMA Flowchart. Literature was included in the study if the following inclusion criteria were met: the study was a randomized controlled trial (RCT); the patients selected for investigation had depression; and the control condition was non-exercise. Acceptable control groups included usual care (standard medical management without structured exercise), waitlist (delayed treatment), no intervention, attention control (e.g., non-exercise health education, relaxation sessions, or light stretching not meeting exercise prescription criteria), or pharmacological treatment (antidepressants alone). Studies where the control group received any structured aerobic, resistance, or mixed exercise training were excluded. Studies where control groups received psychotherapy or cognitive-behavioral therapy without an exercise component were included only if clearly defined as non-exercising.
The study compared exercise interventions to control conditions. Control groups were defined as non-exercising comparators, including usual care, waitlist, no treatment, attention control (e.g., health education, relaxation, or stretching activities that did not meet exercise intensity criteria), or pharmacological treatment alone. Studies were excluded if the control group received any structured exercise intervention or if exercise was compared only to another exercise regimen, without a non-exercising control arm. Studies that did not evaluate the effects of exercise intervention for depression symptoms, studies on depression in patients without exercise or control, and studies with no comparison group were excluded.
Handling of multi-arm studies
For studies with more than two relevant arms (e.g., multiple exercise interventions or multiple comparator groups), the following strategies were applied to avoid double-counting and unit-of-analysis errors as multiple exercise arms vs. single control arm When a study compared two or more different exercise interventions (e.g., aerobic vs. resistance training) to a single control group, the exercise arms were combined into a single pooled exercise group using formulas recommended by the Cochrane Handbook (Chapter 23). This approach avoids double-counting the control group participants while preserving statistical independence; exercise vs. multiple non-exercise comparators. When a study compared exercise to two or more distinct control conditions (e.g., usual care and pharmacotherapy), only the control condition that best aligned with the definition of 'non-exercising control' was selected (priority order: waitlist ≥ usual care ≥ attention control > pharmacotherapy > psychotherapy).
If multiple non-exercising controls were present, the most conservative (least active) comparator to avoid overestimating the effect of exercise was selected; exercise vs. exercise only (no non-exercising control). Studies that compared only different types of exercise (e.g., high-intensity vs. low-intensity) without a non-exercising control arm were excluded from the primary analysis, as they did not meet the comparison criterion of 'exercise vs. control'. Combined interventions for studies, where the intervention arm included exercise plus another active component (e.g., exercise + pharmacotherapy, exercise + mindfulness), the arm was included only if the effect of exercise could be isolated. If the combined arm was compared to the same non-exercise component alone (e.g., exercise + pharmacotherapy vs. pharmacotherapy alone), the difference was interpreted as reflecting the effect of exercise. If no such isolation was possible, the study was excluded. To assess the robustness of the pooling strategy for multi-arm trials, a sensitivity analysis excluding all multi-arm studies was conducted, and the results were compared with those of the primary analysis. Any substantial differences are reported.
Selection of control arm when multiple non‑exercise comparators are present was done in the following manner: if a study compared exercise to two or more distinct control conditions (e.g., usual care and pharmacotherapy), a hierarchical selection rule to avoid arbitrary choices was applied as passive controls (waitlist ≥ usual care ≥ no intervention) prioritized over active controls to better isolate the specific effect of exercise. If only active controls were available, the most conservative comparator (pharmacotherapy ≥ psychotherapy ≥ attention control) to avoid overestimating the exercise effect was selected, and sensitivity analyses were conducted, including all available control arms (by splitting the exercise group across comparators) to test whether the selection rule influenced the pooled estimate. Results of these sensitivity analyses are reported in the Supplementary Materials.
Search strategy
The following electronic databases were systematically searched from inception to November 2025: PubMed, Embase (OVID platform), Cochrane Central Register of Controlled Trials (CENTRAL), MEDLINE (OVID platform), and Google Scholar. The complete search strategies, including MeSH terms, Emtree terms, keywords, Boolean operators, truncation, and field tags, are provided in Table 1. For PubMed and Embase, searches were limited to human studies and the English language. For CENTRAL and MEDLINE, searches were limited to randomized controlled trials. For Google Scholar, the first 500 records sorted by relevance were searched using a simplified search string ('exercise AND depression AND (remission OR medication OR treatment)') due to platform limitations on complex Boolean searching19.
In addition to database searches, the following supplementary searches were conducted: hand-searching of reference lists of all included studies and relevant systematic reviews (snowball searching), citation tracking of key included studies using Google Scholar and Web of Science, hand-searching of key journals (Mental Health and Physical Activity, Journal of Affective Disorders, Depression and Anxiety, and Psychosomatic Medicine) for the years 2020–2025, and searching of clinical trial registries (ClinicalTrials.gov, WHO ICTRP) for unpublished or ongoing trials using the search terms 'exercise' AND 'depression'. The complete search history, including the number of records retrieved at each step for each database, is provided in Supplementary Table 1. The PRISMA 2020 search reporting checklist was followed (Supplementary File 1).
Selection process
All identified records were exported to EndNote 20 for duplicate removal. Following deduplication, two independent reviewers performed a calibration exercise on a random sample of 100 records to ensure consistent application of eligibility criteria. After confirming acceptable agreement (κ ≥ 0.80), the reviewers independently screened titles and abstracts of all remaining records against the prespecified eligibility criteria. Records deemed potentially relevant by either reviewer were sent for full-text review. Full-text articles were obtained and independently assessed by the same two reviewers. Reasons for exclusion at the full-text stage were documented according to PRISMA guidelines (e.g., no exercise intervention, no depression diagnosis, no control group, non-comparative design, duplicate publication). Disagreements at either screening stage were resolved through discussion; if consensus could not be reached, the corresponding author made the final decision. Inter-rater agreement for full-text inclusion was calculated using Cohen’s kappa statistic. Although EndNote 20 was used for initial duplicate removal, additional duplicates that were not detected by the software (e.g., due to slight variations in titles, author names, or journal abbreviations) were identified and removed during manual title/abstract screening by the two independent reviewers.
The screening and selection process is summarized in the PRISMA 2020 flow diagram (Figure 1), which includes the number of records identified, duplicates removed, records screened, reports sought for retrieval, reports assessed for eligibility, and studies included, with specific exclusion reasons at each stage. Google Scholar was searched using the simplified string 'exercise AND depression AND (remission OR medication OR treatment)' with the following settings: sort by relevance, exclude citations, and include patents. The first 500 records were screened. The search was conducted on November 15, 2025, at 10;00 AM GMT, with no personalization enabled, and was repeated on November 16, 2025, to verify consistency.
Data extraction
Two reviewers independently extracted data from each included study using a standardized, pilot-tested data extraction form. The following information was collected: first author name, publication year, country where the study was conducted, sample size (total, exercise group, control group), participant characteristics (age, depression diagnosis or screening method, baseline depression severity), intervention details (exercise mode, intensity, frequency, duration, supervision, format), control group description, depression outcome measures (scale name and range), timing of outcome assessment, and statistical results (means, standard deviations, effect estimates, confidence intervals, p-values)20. Disagreements between reviewers were resolved by discussion or by consulting the corresponding author.
Data items
Data were independently gathered based on a valuation of exercise intervention for depression symptoms when a study produced different results.
Risk of bias assessment
Two authors independently assessed the risk of bias for each included study using the Cochrane RoB 2 tool for randomized controlled trials (version of August 2019). The RoB 2 tool evaluates five domains of bias arising from the randomization process, deviations from intended interventions, missing outcome data, measurement of the outcome, and selection of the reported result. For each domain, reviewers assigned a judgment of ‘low-risk of bias,’ ‘some concerns,’ or ‘high risk of bias.’ An overall risk-of-bias judgment was then derived for each study according to the RoB 2 algorithms ‘low-risk’ (all domains low-risk), ‘some concerns’ (at least one domain with some concerns but no high-risk domains), or ‘high risk’ (any domain rated high risk or multiple domains with some concerns). Disagreements between reviewers were resolved through discussion, with the corresponding author arbitrating when necessary. Studies were not excluded based on risk-of-bias judgments; instead, sensitivity analyses were planned to restrict the meta-analysis to studies with 'low-risk' or 'low-risk + some concerns' overall judgments to examine the robustness of the findings. Inter-rater agreement for overall risk-of-bias judgments was calculated using Cohen’s kappa21.
Subgroup analyses (moderator analyses)
To explore potential sources of heterogeneity and identify moderators of exercise efficacy, a series of a priori subgroup analyses based on clinically and methodologically relevant variables was conducted. Subgroup analyses were performed only for outcomes with at least 10 studies per subgroup (to ensure adequate statistical power). The following subgroup variables were prespecified.
Exercise characteristics
Exercise mode as aerobic (e.g., walking, running, cycling) vs. resistance (e.g., weight training) vs. mixed (aerobic + resistance) vs. other (yoga, dance, mindfulness-based movement).
Exercise intensity as light (e.g., gentle walking, stretching) vs. moderate (e.g., brisk walking, moderate cycling) vs. vigorous (e.g., running, high-intensity interval training) – classified using standard heart rate or perceived exertion criteria reported in each study.
Exercise duration (weeks) as short (≤8 weeks) vs. medium (9–12 weeks) vs. long (≥13 weeks).
Supervision as supervised (exercise sessions led by an instructor, trainer, or therapist) vs. unsupervised (home-based or self-directed with no direct supervision).
Format as group-based vs. individual (home or one-on-one).
Population characteristics
Depression diagnosis was categorized as either major depressive disorder (MDD), diagnosed through a structured clinical interview such as the structured clinical interview for DSM disorders (SCID) or the mini international neuropsychiatric interview (MINI), or elevated depressive symptoms identified using validated screening tools based on established cutoff scores, including the center for epidemiologic studies depression scale (CES-D ≥ 16), Beck depression inventory (BDI ≥ 14), or patient health questionnaire-9 (PHQ-9 ≥ 10). Age group was classified as adults (18–59 years) or older adults (≥60 years). Baseline depression severity was categorized as moderate or severe according to standard cutoff scores for each assessment scale, with moderate depression defined, for example, as a Hamilton Depression Rating Scale (HAM-D) score of 14–18 or a BDI score of 14–19, and severe depression defined as a HAM-D score of ≥19 or a BDI score of ≥20.
Study design characteristics
Control group type was categorized as either passive, including waitlist, no-treatment, or usual care control conditions, or active, including attention control interventions, pharmacotherapy, or psychotherapy. Publication year was classified into three periods—pre-2010, 2010–2019, and 2020–2025—to examine potential temporal trends in study findings. Risk of bias was assessed using the Cochrane Risk of Bias 2 (RoB 2) tool and categorized as low-risk, some concerns, or high risk. Sample size was classified as small (n < 50 total participants), medium (50–99 participants), or large (≥100 participants). Control group type was categorized as passive (waitlist, usual care, no intervention) – these controls do not provide an active therapeutic alternative, so the effect of exercise may be larger due to expectation or attention effects; and active (attention control, pharmacotherapy, psychotherapy) – these controls provide a non‑exercise intervention of credible therapeutic value, yielding a more conservative estimate of exercise’s specific effect. A significantly larger effect for passive vs. active controls (p for interaction < 0.10), as hypothesized, was confirmed in the results (see Table 3).
Statistical methods for subgroup analyses
For each subgroup analysis, studies within each subgroup were pooled using a random-effects model22. To compare subgroups, mixed-effects meta-regression or subgroup interaction tests were used. Specifically, the Q-statistic for between-subgroup differences was calculated, and the associated p-value for interaction is reported. A p-value < 0.10 (rather than 0.05) was considered statistically significant for subgroup differences, acknowledging the low power of interaction tests in meta-analyses. Multiple-comparison adjustment (12 subgroup analyses) was not performed; therefore, subgroup findings should be interpreted as exploratory.
Sensitivity analyses
To assess the robustness of the primary findings, the following prespecified sensitivity analyses were conducted. Model choice as re-analyzed both outcomes using a fixed-effect model (inverse variance method) to compare with the primary random-effects results, exclusion of high-risk-of-bias studies as excluded studies with an overall RoB 2 judgment of ‘high risk of bias’ (n = 17), and re-analyzed the remaining studies (low-risk + some concerns), exclusion of multi-arm trials as excluded studies where multiple exercise arms were combined or where a single control group was used for multiple comparisons (n = 9) and re-analyzed only two-arm trials, exclusion of studies with imputed or assumed data as excluded studies where standard deviations for change scores were assumed (r = 0.50) rather than reported (n = 6), exclusion of pre-CONSORT studies as excluded studies published before 2001 (n = 5) to assess whether improved reporting standards influenced effect estimates, influence analysis (leave-one-out) as performed a leave-one-out meta-analysis, sequentially removing each study and recalculating the pooled effect to identify influential studies (defined as those whose removal changes the point estimate by >10% or moves the confidence interval beyond the original bounds), and long-term follow-up analysis as for studies reporting follow-up data at ≥6 months post-intervention (n = 6), these data were separately pooled to examine the durability of exercise effects, without mixing with primary endpoint (post-intervention) data. All sensitivity analyses were conducted using random-effects models unless otherwise specified. Results of sensitivity analyses are reported in the Results section and supplementary materials. Any deviation from these prespecified analyses is reported transparently.
Extraction of continuous outcome data (depression scores)
For each included study, the mean depression score and its standard deviation (SD) for both exercise and control groups were extracted at the primary endpoint (defined as the time point closest to the end of the exercise intervention). To maximize consistency across studies, the following hierarchical extraction rule was applied: adjusted effect estimates (e.g., ANCOVA-derived between-group differences with 95% CIs or SEs) were prioritized when available, as they account for baseline imbalances and typically provide the least biased estimate of intervention effect. Change-from-baseline scores (post-minus-pre or percentage change) and their SDs were extracted when adjusted effects were not reported. If the SD of change scores was not directly reported, it was calculated using the formula:
(1)
where a conservative correlation coefficient of r = 0.5 was assumed unless the study provided a reported or imputable correlation. Post-intervention scores only (without adjustment or change scores) were extracted when neither adjusted effects nor change scores were available. This was the least preferred option because it does not account for potential baseline differences between groups. Conversion from other statistics, as if means and SDs were not reported directly, available statistics (medians and IQRs, t-statistics, F-statistics, p-values, or 95% CIs) were converted to means and SDs using standard formulas from the Cochrane Handbook (Chapter 6). Studies reporting only medians and ranges were included only if sample sizes were large enough (n ≥ 25 per group) to assume approximate normality, or if the study authors provided means and SDs upon request.
Multiple time points as for studies reporting multiple follow-up time points, the time point closest to the end of the active intervention period (typically 8–12 weeks) was extracted. Long-term follow-up data (≥6 months post-intervention) were extracted separately for a planned secondary analysis of durability, but these data were not pooled with primary endpoint data. Multiple depression scales. When a study reported more than one depression scale (e.g., both HAM-D and BDI), clinician-rated scales (HAM-D, MADRS) were prioritized over self-report scales (BDI, CES-D, PHQ-9, GDS) because clinician-rated scales are considered the gold standard in depression trials. If both were reported, the primary outcome measure as defined by the study authors was extracted. Imputation of missing SDs as if standard deviations were missing and could not be calculated from available statistics; they were imputed by taking the average SD from studies with similar sample sizes and the same depression scale. Sensitivity analyses were conducted, excluding studies with imputed SDs, to assess the impact of this imputation. All data extraction decisions were performed independently by two authors using a standardized data extraction form. Discrepancies were resolved by discussion or by the corresponding author. The extraction rule and any deviations were recorded.
Effect measures
For the dichotomous outcome (depression remission), odds ratios (ORs) with 95% confidence intervals (CIs) were calculated. For the continuous outcome (depression severity scores), included studies were anticipated to use different depression rating scales (e.g., Hamilton Depression Rating Scale, Beck Depression Inventory, PHQ-9, CES-D, Geriatric Depression Scale). Although included studies used different depression rating scales, the standard mean difference (SMD) was reported because all scales had established conversion equations to HAM-D.
Statistical synthesis
For the dichotomous outcome (depression remission), pooled odds ratios (ORs) with 95% confidence intervals (CIs) were calculated. For the continuous outcome (depression severity scores), pooled standardized mean differences (SMDs, Hedges' g) with 95% CIs were calculated because included studies used different depression rating scales. Given the anticipated clinical and methodological diversity across studies (variations in exercise protocols, populations, control conditions, and outcome measures), a random-effects model (DerSimonian and Laird method) was chosen a priori for all primary analyses22. Heterogeneity was quantified using the I2 statistic, interpreted using Cochrane Handbook criteria as next 0–40% (might not be important), 30–60% (moderate heterogeneity), 50–90% (substantial heterogeneity), and 75–100% (considerable heterogeneity). The between-study variance (tau2) was also estimated, and 95% prediction intervals for the pooled effects were calculated where appropriate.
Statistical analysis and heterogeneity
Due to the expected clinical and methodological diversity across included studies (including variations in exercise type, intensity, duration, supervision, population characteristics, depression severity, control conditions, and depression measurement scales), a random-effects model (DerSimonian and Laird method) was chosen a priori for all primary analyses22. The random-effects model assumes that the true effect varies across studies and provides a more conservative estimate with wider confidence intervals, which is appropriate given the anticipated heterogeneity.
Assessment of heterogeneity
Heterogeneity was quantified using the I2 statistic, which represents the percentage of total variation across studies due to true heterogeneity rather than chance. I2 values were interpreted as follows: 0–40% (might not be important); 30–60% (moderate heterogeneity); 50–90% (substantial heterogeneity); 75–100% (considerable heterogeneity), as in Cochrane Handbook criteria. In addition to I2, the between-study variance (tau2) and report 95% prediction intervals for the pooled effect when appropriate were estimated.
Model selection justification
A fixed-effect model was not used for any primary analysis because the studies are not functionally identical (they vary in exercise protocols, populations, and outcome measures), the assumption of a single true effect shared by all studies is implausible in exercise-depression research, and even when I2 appears low (e.g., I2 = 48% for remission), clinical diversity alone justifies a random-effects approach. For completeness, sensitivity analyses using a fixed-effect model were also conducted to examine whether model choice influenced the conclusions; these results are reported in the supplementary materials.
Publication bias assessment
Publication bias and small-study effects were assessed using both visual inspection of funnel plots and Egger's linear regression test. For each outcome, funnel plots were generated showing the effect size (log odds ratio for remission; standardized mean difference for depression scores) against its standard error. For the quantitative assessment, Egger's test regresses the standardized effect estimate against its precision (the inverse of the standard error). A statistically significant intercept (p < 0.10 in meta-analyses with ≥10 studies) suggests the presence of funnel plot asymmetry, which may indicate publication bias or small-study effects. Conversely, p ≥ 0.10 indicates no statistical evidence of such bias.
Prediction intervals
For the primary random-effects meta-analysis, 95% prediction intervals were calculated, which estimate the range within which the true effect of a future study would fall. Wide prediction intervals indicate that the effect of exercise may vary substantially across settings and populations, with important clinical implications. For multi-arm studies in which multiple exercise arms were combined into a single group, the formulas in the Cochrane Handbook were used to calculate combined means, standard deviations, and sample sizes. For studies that used a single control arm across multiple exercise comparisons, double counting was avoided by dividing the control group sample size by the number of exercise arms when calculating effect sizes. All analyses were conducted using Reviewer Manager Version 5.3, with multi-arm handling verified by two authors independently.
Number needed to treat (NNT) calculation
The number needed to treat (NNT) was calculated from the pooled odds ratio (OR) for remission using the formula

where CER (control event rate) was the pooled remission rate across all included studies. NNT was also calculated using the alternative method proposed by Furukawa and Leucht (2011), which converts Cohen's d (from SMD) to NNT using the area under the normal curve. NNT values were calculated for the primary analysis (all studies), for subgroups (low-risk of bias studies, MDD-only studies, supervised exercise studies), and for comparison with published NNT values for psychotherapy and antidepressant medication. An NNT of ≤ 10 was considered clinically meaningful.
Conversion to clinically familiar metrics (BDI and HAM-D)
To aid clinical interpretability, the pooled standardized mean difference (SMD) for depression scores was converted into estimated mean differences on two commonly used depression scales: the Beck Depression Inventory (BDI, range 0–63) and the Hamilton Depression Rating Scale (HAM-D, range 0–52). Conversion used the following formula

where SDpooled, reference was the pooled baseline standard deviation from studies that used the BDI (n = 8 studies) or HAM-D (n = 16 studies), respectively. This approach assumes that the SMD effect size is consistent across scales. The observed mean differences from these studies are also reported as a sensitivity check. A change of ≥3 points on the BDI or HAM-D was considered clinically meaningful per NICE guidelines5.
Comparison with established depression treatments
To contextualize the effect of exercise relative to other first-line depression treatments, the pooled NNT for exercise was compared with published NNT values from recent high-quality meta-analyses. Psychotherapy, as Cuijpers et al. (2020) reported, has an NNT of 2.5 for psychotherapy across all age groups. Antidepressant medication, as Cipriani et al. (2018) reported, has an NNT of 4.3 for medication vs. placebo. These comparisons are descriptive and not based on head-to-head trials. Exercise data were not pooled with psychotherapy or medication data, and differences in study populations, outcome definitions, and time points limit direct comparability.
Assessment of reporting bias (publication bias)
Funnel plots for each outcome (remission and depression score), funnel plots showing the effect size (log odds ratio for remission; standardized mean difference for depression scores) against its standard error were generated. Asymmetry in the funnel plot may indicate publication bias, small-study effects, or heterogeneity. Egger's regression test, also known as Egger's linear regression test, was conducted to formally assess funnel plot asymmetry. Egger's test regresses the standardized effect estimate on its precision (the inverse of the standard error). A statistically significant intercept (p < 0.10 for Egger's test, as recommended by the Cochrane Handbook for outcomes with fewer than 10 studies, or p < 0.05 for larger meta-analyses) suggests the presence of small-study effects or publication bias. Conversely, p ≥ 0.05 (or p ≥ 0.10) indicates no statistically significant evidence of such bias. Interpretation caution as Funnel plot asymmetry is not synonymous with publication bias; it may also arise from true heterogeneity, differences in study quality, or chance, and the results should be interpreted accordingly23.
Certainty of evidence (GRADE assessment)
The overall certainty of evidence for each outcome (remission and depression score) was assessed using the Grading of Recommendations Assessment, Development and Evaluation (GRADE) framework. Two authors independently rated the certainty and resolved disagreements by discussion. The GRADE approach considers five domains that can lower certainty: risk of bias (based on RoB 2 assessments, including the proportion of studies with low-risk, some concerns, or high risk), inconsistency (based on the I2 statistic, prediction intervals, and overlap of confidence intervals), indirectness (based on differences between study populations, interventions, comparators, and outcomes relative to the review question), imprecision (based on optimal information size and whether 95% confidence intervals cross clinically important thresholds, such as an odds ratio of 1.0 for remission or a standardized mean difference of 0.2 for a small effect on depression scores), and publication bias (based on funnel plot asymmetry and Egger's test). Certainty was graded as follows High (Further research is very unlikely to change confidence in the effect estimate); Moderate (Further research is likely to have an important impact on confidence and may change the estimate); Low (Further research is very likely to have an important impact on confidence and is likely to change the estimate); and Very low (Any estimate of effect is very uncertain). Certainty assessments were performed separately for remission (dichotomous outcome) and depression scores (continuous outcome).