Research Article

A Meta-Analysis of Exercise Intervention for Depression Symptoms

97 views

DOI:

10.3791/70951

August 7th, 2026

In This Article

Summary

Exercise significantly increased depression remission rates (OR = 2.40) and reduced symptom severity (SMD = –0.96) in 40 RCTs (2,667 participants). These findings support exercise as a complementary approach, though heterogeneity and bias limit certainty.

Abstract

This meta‑analysis estimated the effect of exercise on depression symptoms. Unlike network meta‑analyses that compare exercise modalities, this pairwise meta‑analysis specifically quantifies the effect of structured exercise versus no exercise, using rigorous control-group definitions and updated methods. A literature search up to November 2025 identified 10,123 records. Forty randomized controlled trials met the inclusion criteria. Pooled effect sizes were calculated using random‑effects models; odds ratios (OR) for remission and standardized mean differences (SMD) for depression scores. Heterogeneity was assessed with I2. Exercise significantly increased remission rates (OR = 2.40; 95% CI, 1.97–2.93; p < 0.001; I2 = 49%) and reduced depression scores (SMD = –0.96; 95% CI, –1.27 to –0.65; p < 0.001; I2 = 87%) compared to control groups. The number needed to treat for one additional remission was 2.1. The pooled SMD corresponded to clinically meaningful reductions. Subgroup analyses showed larger effects for supervised exercise and passive controls (e.g., waitlist) versus active comparators (e.g., pharmacotherapy). However, findings have major limitations; only 12/40 studies (30%) had low-risk of bias (RoB 2); most had short follow‑up (≤16 weeks); control conditions varied widely (waitlist to pharmacotherapy); and very high heterogeneity for depression scores (I2 = 87%) persisted despite using SMD. The 95% prediction interval for SMD crossed zero, indicating exercise may not benefit all future populations. Grading of Recommendations Assessment, Development and Evaluation (GRADE) was moderate for remission but low for depression scores. Despite these limitations, exercise was associated with significantly higher remission and lower depression scores. These results support exercise as a complementary approach for depression, particularly for achieving remission. However, high‑quality, large‑scale RCTs with standardized outcomes and longer follow‑up are needed before clinical recommendations can be made.

Introduction

Depression is a common and incapacitating mental illness linked to higher mortality, medical comorbidity, and lower quality of life1. More than 300 million people, or around 4.4% of the global population, suffer from depression2. Psychotherapy and antidepressant medication (or a mix of both) are currently advised therapies3. However, psychotherapy is usually expensive and only produces remission rates of 50%4. Antidepressant medication frequently causes withdrawal symptoms, relapses, and side effects5. Notably, around two-thirds of adults who suffer from depression do not obtain proper care6. Depression that goes untreated frequently worsens and can result in comorbidities, which raises the expenses to society even more7. Both the national institute for health and care excellence (NICE) and the WHO advocate exercise as a supplemental treatment for depression8,9.

Results from several meta-analyses examining the antidepressant impact of exercise on individuals with depression provided evidence in support of these suggestions10,11,12. Nevertheless, some of these meta-analyses8 indicated significant effects of exercise, whereas others identified moderate to weak effects11,12. Methodological and conceptual disagreements over inclusion criteria and analysis techniques are the cause of these conflicting findings. For instance, some research excluded studies that assessed the existence of depression using established screening tools and instead concentrated on those with a major depressive disorder diagnosis8. For research published between 2003 and 2019, others examined the impact of exercise either by itself or in conjunction with pharmaceutical treatments as a treatment for depression11,12. Additionally, trials where patients in the control groups also received exercise interventions13 were included in certain reviews. Given that even mild exercise might have antidepressant benefits, this raises the possibility of bias14.

Crucially, a number of evaluations have expressed concern that, when limited to "low-risk of bias" randomized controlled trials (RCTs), exercise had no discernible impact8,13. As a result, current meta-analyses have not produced strong enough data to support the use of exercise as an evidence-based, successful depression therapy choice by doctors worldwide. These methodological flaws were addressed by one meta-analysis11, which only included studies that compared exercise to non-active controls and concentrated on studies that included samples with major depressive disorder diagnosed using diagnostic tools and samples with depression using cut-offs on validated screening instruments. Trials comparing various exercise regimens were not included by the authors. However, a significant number of papers have been published recently, necessitating an updated meta-analysis of exercise's antidepressant effects that addresses the limitations of earlier reviews.

This demonstrates the necessity of quick, accessible alternative treatment options. Recent systematic reviews have confirmed that structured physical activity and exercise programs consistently improve depressive symptoms and related mental health outcomes, supporting exercise as a potentially scalable adjunctive or complementary approach15,16. The manuscript includes multiple forms of exercise, including aerobic exercise, resistance training, mixed exercise, yoga, tai chi, dance, and other mind-body approaches. However, the discussion still tends to describe "exercise" as a single intervention category15. Despite the growing evidence base for exercise as a treatment for depression, several unresolved issues justify an updated meta-analysis. First, prior meta-analyses have produced conflicting conclusions. These discrepancies arise from methodological differences, including the inclusion of studies with active control groups (e.g., stretching, relaxation, or psychotherapy), which may attenuate effect sizes; the inclusion of studies in which control groups also received exercise interventions; and the use of different risk-of-bias assessment tools. Second, the most recent comprehensive meta-analyses in this area were published in 2016–201910,11,12. Since 2019, at least 23 new randomized controlled trials have been published (as identified in this search), substantially increasing the available evidence base. An updated synthesis is therefore necessary to incorporate this new data and determine whether the cumulative evidence supports or refutes earlier conclusions.

Third, prior meta-analyses have not consistently addressed key methodological issues that affect clinical interpretation. These include the use of appropriate effect metrics for continuous outcomes measured on different depression scales (standardized mean difference rather than mean difference), the handling of heterogeneous control conditions (passive vs. active comparators), the inclusion of studies where depression was a secondary outcome in non-psychiatric populations, the assessment of publication bias in the context of high heterogeneity, and the calculation of prediction intervals to communicate the range of expected effects in future studies. Fourth, no prior meta-analysis has systematically examined moderators of exercise efficacy across the full range of potential effect modifiers, including exercise mode (aerobic, resistance, mixed, yoga), intensity (light, moderate, vigorous), duration, supervision, format (group vs. individual), control type, depression diagnosis (MDD vs. screening-based), age group, and risk of bias. Understanding these moderators is essential for prescribing exercise.

This meta-analysis addresses these gaps through the following innovations. First, an updated evidence base includes studies published through November 2025, capturing 23 new trials not included in prior meta-analyses. Next, a methodologically rigorous inclusion criterion was used, where only studies that compared exercise to non-exercising control groups (excluding trials in which control groups received stretching, relaxation, or active psychological interventions) were included to isolate the specific effect of exercise. Third was the appropriate effect metric, with standardized mean difference (SMD) rather than mean difference, for continuous depression outcomes, acknowledging that studies used different depression scales. Fourth was comprehensive moderator analyses, in which prespecified subgroup analyses of exercise characteristics, population characteristics, and study design features were conducted to identify conditions under which exercise is most effective. Next was the current risk-of-bias assessment in which the updated RoB 2 tool (2019 version) was used rather than the older RoB 1 tool employed in prior meta-analyses. Another one was clinical interpretability in which the number needed to treat (NNT) was calculated, SMD was converted to clinically familiar metrics (BDI and HAM-D points), and prediction intervals were reported to communicate the expected range of effects in future applications. The final one was the certainty of evidence, in which the GRADE (Grading of Recommendations Assessment, Development and Evaluation) was applied to rate the certainty of evidence for each outcome, a rating that has not been consistently reported in prior exercise-depression meta-analyses.

Although recent network meta‑analyses (NMAs) and umbrella reviews have synthesized evidence on exercise for depression, none directly answer the foundational pairwise question, "Does any structured exercise reduce depressive symptoms compared to no exercise?" NMAs pool across heterogeneous control conditions (waitlist, pharmacotherapy, psychotherapy) to compare modalities, which obscures the specific effect of exercise versus true inactivity. Umbrella reviews aggregate prior meta‑analyses that used different inclusion criteria and risk‑of‑bias tools, leaving methodological heterogeneity unresolved. This pairwise meta‑analysis fills this specific gap by providing a single, up‑to‑date, internally consistent estimate of exercise versus non‑exercising control using prespecified eligibility, RoB 2, GRADE, prediction intervals, and NNT – including 23 new trials not yet synthesized in this framework. This foundational estimate complements NMAs by answering whether exercise works before asking which exercise works best.

Exercise is a behavioral intervention, not a medication. Throughout this manuscript, exercise is compared to pharmacotherapy and psychotherapy using metrics such as number needed to treat (NNT) and remission rates. This “exercise as medication” framing is useful for clinical communication and guideline development because it allows direct comparison of effect sizes across treatment modalities using common metrics. However, important distinctions must be acknowledged. Exercise is a complex behavioral intervention that requires active patient participation, motivation, and often supervision or environmental support. Unlike a pill, exercise has no direct biological target, no standardized dose–response curve across individuals, and adherence is heavily influenced by psychosocial factors such as depression-related anhedonia and fatigue. Moreover, exercise does not carry the same side-effect profile (e.g., no withdrawal syndromes or metabolic disturbances) but does pose injury risk and logistical barriers. Therefore, while comparative effect estimates are presented, the intention is not to suggest that exercise is interchangeable with medication; rather, exercise should be viewed as a complementary, patient‑centered option with distinct implementation requirements.

This meta-analysis aimed to compare exercise with non-exercising control groups in order to update the current evidence on the effects of exercise in lowering depressive symptoms in individuals with clinically elevated levels of depression, including major depressive disorder and dysthymia. The existence of publication bias and possible moderators of exercise's antidepressant effects were also examined. Hence, the meta-analysis evaluated exercise intervention for depression symptoms.

Protocol

Eligibility criteria

Studies were included if they met all of the following criteria, organized according to the PICOS framework (Population, Intervention, Comparator, Outcomes, Study design)17. All platforms and materials used in this study are listed in the Table of Materials. All database searches were performed in November 2025. No additional proprietary software (e.g., SAS, SPSS, R) was used beyond RevMan 5.3 and EndNote 20. The GRADE assessment was performed manually using the GRADEpro GDT online tool; the version is recorded as of the access date.

Population (P)

Participants had clinically significant depressive symptoms, defined as either a formal diagnosis of major depressive disorder (MDD) or dysthymia according to DSM criteria (any version) or ICD criteria, confirmed by structured clinical interview (e.g., SCID, MINI, CIDI); or a score above the validated cutoff on a standardized depression screening instrument (e.g., Beck Depression Inventory ≥ 14, Hamilton Depression Rating Scale ≥ 14, CES-D ≥ 16, PHQ-9 ≥ 10, Geriatric Depression Scale ≥ 11). Participants could be of any age (children, adolescents, adults, older adults)18. Studies that enrolled participants with comorbid medical or psychiatric conditions (e.g., cardiovascular disease, diabetes, anxiety disorders, substance use disorders) were included only if the primary analysis reported results separately for the depression subgroup or if ≥80% of participants met depression criteria. Studies where depression was a secondary outcome in a non-depressed population (e.g., heart failure patients without depression screening) were excluded. There was no restriction on baseline depression severity (mild, moderate, severe), but severity was extracted for subgroup analyses.

Intervention (I)

The intervention consisted of structured, planned physical exercise, defined as any form of aerobic, resistance, mixed (aerobic + resistance), or alternative exercise (yoga, tai chi, dance) delivered for at least 2 weeks. Exercise interventions could be supervised (by a trainer, therapist, or instructor) or unsupervised (home-based, self-directed). Studies where exercise was combined with another active intervention (e.g., exercise + pharmacotherapy, exercise + mindfulness) were included only if the effect of exercise could be isolated (e.g., exercise + pharmacotherapy vs. pharmacotherapy alone). Studies where the intervention consisted only of physical activity advice (without structured exercise prescription) or lifestyle counseling were excluded.

Comparator (C)

The control group received no structured exercise intervention. Acceptable control conditions were categorized as passive controls (no active therapeutic intervention) e.g. waitlist (delayed treatment), usual care (standard medical management without structured exercise), or no intervention; and active controls (non‑exercise therapeutic interventions) e.g. attention control (e.g., non‑exercise health education, relaxation sessions, or light stretching not meeting exercise intensity criteria), pharmacotherapy (antidepressant medication alone), or psychotherapy (psychological therapy without exercise). Studies in which the control group received any form of structured exercise (including minimal exercise or stretching that met the exercise intensity criteria) were excluded. Note on discrepancy, although the innovation section stated that studies with active psychological interventions or stretching/relaxation controls were excluded, subsequently, the eligibility criteria to include active comparators (pharmacotherapy, psychotherapy, attention control) to reflect the real‑world clinical question of how exercise compares to other established treatments were broadened. This discrepancy is addressed in the discussion. All analyses stratify by control type (passive vs. active) to examine the influence of control type on effect sizes.

Outcomes (O)

The study reported at least one depression-specific outcome measure: depression remission (dichotomous, defined by study authors as no longer meeting diagnostic criteria or scoring below the cutoff on a validated scale); OR (b) change in continuous depression severity score measured on a validated scale (e.g., HAM-D, BDI, CES-D, PHQ-9, GDS, MADRS). Studies reporting only non-specific mood measures (e.g., “well-being,” “emotional state” without a validated depression scale) were excluded.

Study design (S)

Only randomized controlled trials (RCTs) were included. Prospective cohort studies, case-control studies, retrospective studies, case series, and case reports were excluded because they are at high risk of confounding and selection bias when estimating treatment effects. Studies could be parallel-group or crossover design (only the first phase extracted for crossover trials). Studies were required to have at least 10 participants per arm at baseline to minimize small-study effects. No language restrictions were applied, but non-English studies were included only if a full translation was available.

Exclusion criteria

Studies where depression was not the primary outcome or where the population was not selected for depression (e.g., general medical populations with depression as a secondary measure) were excluded. Studies in which the exercise intervention was delivered as part of a multi-component lifestyle intervention, without an isolated exercise effect, were excluded. Studies where the control group received any form of structured exercise were excluded. Studies with insufficient data to calculate effect sizes (means, standard deviations, or change scores not reported and not obtainable from authors) were excluded. Duplicate publications of the same study cohort (only the most complete or most recent report was included) were excluded. Conference abstracts, dissertations, and unpublished manuscripts without full-text availability were excluded.

Information sources

The full study is shown in Figure 1 as a PRISMA Flowchart. Literature was included in the study if the following inclusion criteria were met: the study was a randomized controlled trial (RCT); the patients selected for investigation had depression; and the control condition was non-exercise. Acceptable control groups included usual care (standard medical management without structured exercise), waitlist (delayed treatment), no intervention, attention control (e.g., non-exercise health education, relaxation sessions, or light stretching not meeting exercise prescription criteria), or pharmacological treatment (antidepressants alone). Studies where the control group received any structured aerobic, resistance, or mixed exercise training were excluded. Studies where control groups received psychotherapy or cognitive-behavioral therapy without an exercise component were included only if clearly defined as non-exercising.

The study compared exercise interventions to control conditions. Control groups were defined as non-exercising comparators, including usual care, waitlist, no treatment, attention control (e.g., health education, relaxation, or stretching activities that did not meet exercise intensity criteria), or pharmacological treatment alone. Studies were excluded if the control group received any structured exercise intervention or if exercise was compared only to another exercise regimen, without a non-exercising control arm. Studies that did not evaluate the effects of exercise intervention for depression symptoms, studies on depression in patients without exercise or control, and studies with no comparison group were excluded.

Handling of multi-arm studies
For studies with more than two relevant arms (e.g., multiple exercise interventions or multiple comparator groups), the following strategies were applied to avoid double-counting and unit-of-analysis errors as multiple exercise arms vs. single control arm When a study compared two or more different exercise interventions (e.g., aerobic vs. resistance training) to a single control group, the exercise arms were combined into a single pooled exercise group using formulas recommended by the Cochrane Handbook (Chapter 23). This approach avoids double-counting the control group participants while preserving statistical independence; exercise vs. multiple non-exercise comparators. When a study compared exercise to two or more distinct control conditions (e.g., usual care and pharmacotherapy), only the control condition that best aligned with the definition of 'non-exercising control' was selected (priority order: waitlist ≥ usual care ≥ attention control > pharmacotherapy > psychotherapy).

If multiple non-exercising controls were present, the most conservative (least active) comparator to avoid overestimating the effect of exercise was selected; exercise vs. exercise only (no non-exercising control). Studies that compared only different types of exercise (e.g., high-intensity vs. low-intensity) without a non-exercising control arm were excluded from the primary analysis, as they did not meet the comparison criterion of 'exercise vs. control'. Combined interventions for studies, where the intervention arm included exercise plus another active component (e.g., exercise + pharmacotherapy, exercise + mindfulness), the arm was included only if the effect of exercise could be isolated. If the combined arm was compared to the same non-exercise component alone (e.g., exercise + pharmacotherapy vs. pharmacotherapy alone), the difference was interpreted as reflecting the effect of exercise. If no such isolation was possible, the study was excluded. To assess the robustness of the pooling strategy for multi-arm trials, a sensitivity analysis excluding all multi-arm studies was conducted, and the results were compared with those of the primary analysis. Any substantial differences are reported.

Selection of control arm when multiple non‑exercise comparators are present was done in the following manner: if a study compared exercise to two or more distinct control conditions (e.g., usual care and pharmacotherapy), a hierarchical selection rule to avoid arbitrary choices was applied as passive controls (waitlist ≥ usual care ≥ no intervention) prioritized over active controls to better isolate the specific effect of exercise. If only active controls were available, the most conservative comparator (pharmacotherapy ≥ psychotherapy ≥ attention control) to avoid overestimating the exercise effect was selected, and sensitivity analyses were conducted, including all available control arms (by splitting the exercise group across comparators) to test whether the selection rule influenced the pooled estimate. Results of these sensitivity analyses are reported in the Supplementary Materials.

Search strategy
The following electronic databases were systematically searched from inception to November 2025: PubMed, Embase (OVID platform), Cochrane Central Register of Controlled Trials (CENTRAL), MEDLINE (OVID platform), and Google Scholar. The complete search strategies, including MeSH terms, Emtree terms, keywords, Boolean operators, truncation, and field tags, are provided in Table 1. For PubMed and Embase, searches were limited to human studies and the English language. For CENTRAL and MEDLINE, searches were limited to randomized controlled trials. For Google Scholar, the first 500 records sorted by relevance were searched using a simplified search string ('exercise AND depression AND (remission OR medication OR treatment)') due to platform limitations on complex Boolean searching19.

In addition to database searches, the following supplementary searches were conducted: hand-searching of reference lists of all included studies and relevant systematic reviews (snowball searching), citation tracking of key included studies using Google Scholar and Web of Science, hand-searching of key journals (Mental Health and Physical Activity, Journal of Affective Disorders, Depression and Anxiety, and Psychosomatic Medicine) for the years 2020–2025, and searching of clinical trial registries (ClinicalTrials.gov, WHO ICTRP) for unpublished or ongoing trials using the search terms 'exercise' AND 'depression'. The complete search history, including the number of records retrieved at each step for each database, is provided in Supplementary Table 1. The PRISMA 2020 search reporting checklist was followed (Supplementary File 1).

Selection process

All identified records were exported to EndNote 20 for duplicate removal. Following deduplication, two independent reviewers performed a calibration exercise on a random sample of 100 records to ensure consistent application of eligibility criteria. After confirming acceptable agreement (κ ≥ 0.80), the reviewers independently screened titles and abstracts of all remaining records against the prespecified eligibility criteria. Records deemed potentially relevant by either reviewer were sent for full-text review. Full-text articles were obtained and independently assessed by the same two reviewers. Reasons for exclusion at the full-text stage were documented according to PRISMA guidelines (e.g., no exercise intervention, no depression diagnosis, no control group, non-comparative design, duplicate publication). Disagreements at either screening stage were resolved through discussion; if consensus could not be reached, the corresponding author made the final decision. Inter-rater agreement for full-text inclusion was calculated using Cohen’s kappa statistic. Although EndNote 20 was used for initial duplicate removal, additional duplicates that were not detected by the software (e.g., due to slight variations in titles, author names, or journal abbreviations) were identified and removed during manual title/abstract screening by the two independent reviewers.

The screening and selection process is summarized in the PRISMA 2020 flow diagram (Figure 1), which includes the number of records identified, duplicates removed, records screened, reports sought for retrieval, reports assessed for eligibility, and studies included, with specific exclusion reasons at each stage. Google Scholar was searched using the simplified string 'exercise AND depression AND (remission OR medication OR treatment)' with the following settings: sort by relevance, exclude citations, and include patents. The first 500 records were screened. The search was conducted on November 15, 2025, at 10;00 AM GMT, with no personalization enabled, and was repeated on November 16, 2025, to verify consistency.

Data extraction

Two reviewers independently extracted data from each included study using a standardized, pilot-tested data extraction form. The following information was collected: first author name, publication year, country where the study was conducted, sample size (total, exercise group, control group), participant characteristics (age, depression diagnosis or screening method, baseline depression severity), intervention details (exercise mode, intensity, frequency, duration, supervision, format), control group description, depression outcome measures (scale name and range), timing of outcome assessment, and statistical results (means, standard deviations, effect estimates, confidence intervals, p-values)20. Disagreements between reviewers were resolved by discussion or by consulting the corresponding author.

Data items

Data were independently gathered based on a valuation of exercise intervention for depression symptoms when a study produced different results.

Risk of bias assessment

Two authors independently assessed the risk of bias for each included study using the Cochrane RoB 2 tool for randomized controlled trials (version of August 2019). The RoB 2 tool evaluates five domains of bias arising from the randomization process, deviations from intended interventions, missing outcome data, measurement of the outcome, and selection of the reported result. For each domain, reviewers assigned a judgment of ‘low-risk of bias,’ ‘some concerns,’ or ‘high risk of bias.’ An overall risk-of-bias judgment was then derived for each study according to the RoB 2 algorithms ‘low-risk’ (all domains low-risk), ‘some concerns’ (at least one domain with some concerns but no high-risk domains), or ‘high risk’ (any domain rated high risk or multiple domains with some concerns). Disagreements between reviewers were resolved through discussion, with the corresponding author arbitrating when necessary. Studies were not excluded based on risk-of-bias judgments; instead, sensitivity analyses were planned to restrict the meta-analysis to studies with 'low-risk' or 'low-risk + some concerns' overall judgments to examine the robustness of the findings. Inter-rater agreement for overall risk-of-bias judgments was calculated using Cohen’s kappa21.

Subgroup analyses (moderator analyses)

To explore potential sources of heterogeneity and identify moderators of exercise efficacy, a series of a priori subgroup analyses based on clinically and methodologically relevant variables was conducted. Subgroup analyses were performed only for outcomes with at least 10 studies per subgroup (to ensure adequate statistical power). The following subgroup variables were prespecified.

Exercise characteristics

Exercise mode as aerobic (e.g., walking, running, cycling) vs. resistance (e.g., weight training) vs. mixed (aerobic + resistance) vs. other (yoga, dance, mindfulness-based movement).

Exercise intensity as light (e.g., gentle walking, stretching) vs. moderate (e.g., brisk walking, moderate cycling) vs. vigorous (e.g., running, high-intensity interval training) – classified using standard heart rate or perceived exertion criteria reported in each study.

Exercise duration (weeks) as short (≤8 weeks) vs. medium (9–12 weeks) vs. long (≥13 weeks).

Supervision as supervised (exercise sessions led by an instructor, trainer, or therapist) vs. unsupervised (home-based or self-directed with no direct supervision).

Format as group-based vs. individual (home or one-on-one).

Population characteristics

Depression diagnosis was categorized as either major depressive disorder (MDD), diagnosed through a structured clinical interview such as the structured clinical interview for DSM disorders (SCID) or the mini international neuropsychiatric interview (MINI), or elevated depressive symptoms identified using validated screening tools based on established cutoff scores, including the center for epidemiologic studies depression scale (CES-D ≥ 16), Beck depression inventory (BDI ≥ 14), or patient health questionnaire-9 (PHQ-9 ≥ 10). Age group was classified as adults (18–59 years) or older adults (≥60 years). Baseline depression severity was categorized as moderate or severe according to standard cutoff scores for each assessment scale, with moderate depression defined, for example, as a Hamilton Depression Rating Scale (HAM-D) score of 14–18 or a BDI score of 14–19, and severe depression defined as a HAM-D score of ≥19 or a BDI score of ≥20.

Study design characteristics

Control group type was categorized as either passive, including waitlist, no-treatment, or usual care control conditions, or active, including attention control interventions, pharmacotherapy, or psychotherapy. Publication year was classified into three periods—pre-2010, 2010–2019, and 2020–2025—to examine potential temporal trends in study findings. Risk of bias was assessed using the Cochrane Risk of Bias 2 (RoB 2) tool and categorized as low-risk, some concerns, or high risk. Sample size was classified as small (n < 50 total participants), medium (50–99 participants), or large (≥100 participants). Control group type was categorized as passive (waitlist, usual care, no intervention) – these controls do not provide an active therapeutic alternative, so the effect of exercise may be larger due to expectation or attention effects; and active (attention control, pharmacotherapy, psychotherapy) – these controls provide a non‑exercise intervention of credible therapeutic value, yielding a more conservative estimate of exercise’s specific effect. A significantly larger effect for passive vs. active controls (p for interaction < 0.10), as hypothesized, was confirmed in the results (see Table 3).

Statistical methods for subgroup analyses

For each subgroup analysis, studies within each subgroup were pooled using a random-effects model22. To compare subgroups, mixed-effects meta-regression or subgroup interaction tests were used. Specifically, the Q-statistic for between-subgroup differences was calculated, and the associated p-value for interaction is reported. A p-value < 0.10 (rather than 0.05) was considered statistically significant for subgroup differences, acknowledging the low power of interaction tests in meta-analyses. Multiple-comparison adjustment (12 subgroup analyses) was not performed; therefore, subgroup findings should be interpreted as exploratory.

Sensitivity analyses

To assess the robustness of the primary findings, the following prespecified sensitivity analyses were conducted. Model choice as re-analyzed both outcomes using a fixed-effect model (inverse variance method) to compare with the primary random-effects results, exclusion of high-risk-of-bias studies as excluded studies with an overall RoB 2 judgment of ‘high risk of bias’ (n = 17), and re-analyzed the remaining studies (low-risk + some concerns), exclusion of multi-arm trials as excluded studies where multiple exercise arms were combined or where a single control group was used for multiple comparisons (n = 9) and re-analyzed only two-arm trials, exclusion of studies with imputed or assumed data as excluded studies where standard deviations for change scores were assumed (r = 0.50) rather than reported (n = 6), exclusion of pre-CONSORT studies as excluded studies published before 2001 (n = 5) to assess whether improved reporting standards influenced effect estimates, influence analysis (leave-one-out) as performed a leave-one-out meta-analysis, sequentially removing each study and recalculating the pooled effect to identify influential studies (defined as those whose removal changes the point estimate by >10% or moves the confidence interval beyond the original bounds), and long-term follow-up analysis as for studies reporting follow-up data at ≥6 months post-intervention (n = 6), these data were separately pooled to examine the durability of exercise effects, without mixing with primary endpoint (post-intervention) data. All sensitivity analyses were conducted using random-effects models unless otherwise specified. Results of sensitivity analyses are reported in the Results section and supplementary materials. Any deviation from these prespecified analyses is reported transparently.

Extraction of continuous outcome data (depression scores)

For each included study, the mean depression score and its standard deviation (SD) for both exercise and control groups were extracted at the primary endpoint (defined as the time point closest to the end of the exercise intervention). To maximize consistency across studies, the following hierarchical extraction rule was applied: adjusted effect estimates (e.g., ANCOVA-derived between-group differences with 95% CIs or SEs) were prioritized when available, as they account for baseline imbalances and typically provide the least biased estimate of intervention effect. Change-from-baseline scores (post-minus-pre or percentage change) and their SDs were extracted when adjusted effects were not reported. If the SD of change scores was not directly reported, it was calculated using the formula:

Standard deviation change formula, SD_change = √(SD^2_baseline + SD^2_post - 2rSD_baselineSD_post).   (1)

where a conservative correlation coefficient of r = 0.5 was assumed unless the study provided a reported or imputable correlation. Post-intervention scores only (without adjustment or change scores) were extracted when neither adjusted effects nor change scores were available. This was the least preferred option because it does not account for potential baseline differences between groups. Conversion from other statistics, as if means and SDs were not reported directly, available statistics (medians and IQRs, t-statistics, F-statistics, p-values, or 95% CIs) were converted to means and SDs using standard formulas from the Cochrane Handbook (Chapter 6). Studies reporting only medians and ranges were included only if sample sizes were large enough (n ≥ 25 per group) to assume approximate normality, or if the study authors provided means and SDs upon request.

Multiple time points as for studies reporting multiple follow-up time points, the time point closest to the end of the active intervention period (typically 8–12 weeks) was extracted. Long-term follow-up data (≥6 months post-intervention) were extracted separately for a planned secondary analysis of durability, but these data were not pooled with primary endpoint data. Multiple depression scales. When a study reported more than one depression scale (e.g., both HAM-D and BDI), clinician-rated scales (HAM-D, MADRS) were prioritized over self-report scales (BDI, CES-D, PHQ-9, GDS) because clinician-rated scales are considered the gold standard in depression trials. If both were reported, the primary outcome measure as defined by the study authors was extracted. Imputation of missing SDs as if standard deviations were missing and could not be calculated from available statistics; they were imputed by taking the average SD from studies with similar sample sizes and the same depression scale. Sensitivity analyses were conducted, excluding studies with imputed SDs, to assess the impact of this imputation. All data extraction decisions were performed independently by two authors using a standardized data extraction form. Discrepancies were resolved by discussion or by the corresponding author. The extraction rule and any deviations were recorded.

Effect measures

For the dichotomous outcome (depression remission), odds ratios (ORs) with 95% confidence intervals (CIs) were calculated. For the continuous outcome (depression severity scores), included studies were anticipated to use different depression rating scales (e.g., Hamilton Depression Rating Scale, Beck Depression Inventory, PHQ-9, CES-D, Geriatric Depression Scale). Although included studies used different depression rating scales, the standard mean difference (SMD) was reported because all scales had established conversion equations to HAM-D.

Statistical synthesis

For the dichotomous outcome (depression remission), pooled odds ratios (ORs) with 95% confidence intervals (CIs) were calculated. For the continuous outcome (depression severity scores), pooled standardized mean differences (SMDs, Hedges' g) with 95% CIs were calculated because included studies used different depression rating scales. Given the anticipated clinical and methodological diversity across studies (variations in exercise protocols, populations, control conditions, and outcome measures), a random-effects model (DerSimonian and Laird method) was chosen a priori for all primary analyses22. Heterogeneity was quantified using the I2 statistic, interpreted using Cochrane Handbook criteria as next 0–40% (might not be important), 30–60% (moderate heterogeneity), 50–90% (substantial heterogeneity), and 75–100% (considerable heterogeneity). The between-study variance (tau2) was also estimated, and 95% prediction intervals for the pooled effects were calculated where appropriate.

Statistical analysis and heterogeneity

Due to the expected clinical and methodological diversity across included studies (including variations in exercise type, intensity, duration, supervision, population characteristics, depression severity, control conditions, and depression measurement scales), a random-effects model (DerSimonian and Laird method) was chosen a priori for all primary analyses22. The random-effects model assumes that the true effect varies across studies and provides a more conservative estimate with wider confidence intervals, which is appropriate given the anticipated heterogeneity.

Assessment of heterogeneity

Heterogeneity was quantified using the I2 statistic, which represents the percentage of total variation across studies due to true heterogeneity rather than chance. I2 values were interpreted as follows: 0–40% (might not be important); 30–60% (moderate heterogeneity); 50–90% (substantial heterogeneity); 75–100% (considerable heterogeneity), as in Cochrane Handbook criteria. In addition to I2, the between-study variance (tau2) and report 95% prediction intervals for the pooled effect when appropriate were estimated.

Model selection justification

A fixed-effect model was not used for any primary analysis because the studies are not functionally identical (they vary in exercise protocols, populations, and outcome measures), the assumption of a single true effect shared by all studies is implausible in exercise-depression research, and even when I2 appears low (e.g., I2 = 48% for remission), clinical diversity alone justifies a random-effects approach. For completeness, sensitivity analyses using a fixed-effect model were also conducted to examine whether model choice influenced the conclusions; these results are reported in the supplementary materials.

Publication bias assessment

Publication bias and small-study effects were assessed using both visual inspection of funnel plots and Egger's linear regression test. For each outcome, funnel plots were generated showing the effect size (log odds ratio for remission; standardized mean difference for depression scores) against its standard error. For the quantitative assessment, Egger's test regresses the standardized effect estimate against its precision (the inverse of the standard error). A statistically significant intercept (p < 0.10 in meta-analyses with ≥10 studies) suggests the presence of funnel plot asymmetry, which may indicate publication bias or small-study effects. Conversely, p ≥ 0.10 indicates no statistical evidence of such bias.

Prediction intervals

For the primary random-effects meta-analysis, 95% prediction intervals were calculated, which estimate the range within which the true effect of a future study would fall. Wide prediction intervals indicate that the effect of exercise may vary substantially across settings and populations, with important clinical implications. For multi-arm studies in which multiple exercise arms were combined into a single group, the formulas in the Cochrane Handbook were used to calculate combined means, standard deviations, and sample sizes. For studies that used a single control arm across multiple exercise comparisons, double counting was avoided by dividing the control group sample size by the number of exercise arms when calculating effect sizes. All analyses were conducted using Reviewer Manager Version 5.3, with multi-arm handling verified by two authors independently.

Number needed to treat (NNT) calculation

The number needed to treat (NNT) was calculated from the pooled odds ratio (OR) for remission using the formula

NNT formula: 1/(CER×(1−OR^-1)); equation for calculating number needed to treat.

where CER (control event rate) was the pooled remission rate across all included studies. NNT was also calculated using the alternative method proposed by Furukawa and Leucht (2011), which converts Cohen's d (from SMD) to NNT using the area under the normal curve. NNT values were calculated for the primary analysis (all studies), for subgroups (low-risk of bias studies, MDD-only studies, supervised exercise studies), and for comparison with published NNT values for psychotherapy and antidepressant medication. An NNT of ≤ 10 was considered clinically meaningful.

Conversion to clinically familiar metrics (BDI and HAM-D)

To aid clinical interpretability, the pooled standardized mean difference (SMD) for depression scores was converted into estimated mean differences on two commonly used depression scales: the Beck Depression Inventory (BDI, range 0–63) and the Hamilton Depression Rating Scale (HAM-D, range 0–52). Conversion used the following formula

Mathematical formula for estimated mean difference; SMD x SD_pooled; reference value equation.

where SDpooled, reference was the pooled baseline standard deviation from studies that used the BDI (n = 8 studies) or HAM-D (n = 16 studies), respectively. This approach assumes that the SMD effect size is consistent across scales. The observed mean differences from these studies are also reported as a sensitivity check. A change of ≥3 points on the BDI or HAM-D was considered clinically meaningful per NICE guidelines5.

Comparison with established depression treatments

To contextualize the effect of exercise relative to other first-line depression treatments, the pooled NNT for exercise was compared with published NNT values from recent high-quality meta-analyses. Psychotherapy, as Cuijpers et al. (2020) reported, has an NNT of 2.5 for psychotherapy across all age groups. Antidepressant medication, as Cipriani et al. (2018) reported, has an NNT of 4.3 for medication vs. placebo. These comparisons are descriptive and not based on head-to-head trials. Exercise data were not pooled with psychotherapy or medication data, and differences in study populations, outcome definitions, and time points limit direct comparability.

Assessment of reporting bias (publication bias)

Funnel plots for each outcome (remission and depression score), funnel plots showing the effect size (log odds ratio for remission; standardized mean difference for depression scores) against its standard error were generated. Asymmetry in the funnel plot may indicate publication bias, small-study effects, or heterogeneity. Egger's regression test, also known as Egger's linear regression test, was conducted to formally assess funnel plot asymmetry. Egger's test regresses the standardized effect estimate on its precision (the inverse of the standard error). A statistically significant intercept (p < 0.10 for Egger's test, as recommended by the Cochrane Handbook for outcomes with fewer than 10 studies, or p < 0.05 for larger meta-analyses) suggests the presence of small-study effects or publication bias. Conversely, p ≥ 0.05 (or p ≥ 0.10) indicates no statistically significant evidence of such bias. Interpretation caution as Funnel plot asymmetry is not synonymous with publication bias; it may also arise from true heterogeneity, differences in study quality, or chance, and the results should be interpreted accordingly23.

Certainty of evidence (GRADE assessment)

The overall certainty of evidence for each outcome (remission and depression score) was assessed using the Grading of Recommendations Assessment, Development and Evaluation (GRADE) framework. Two authors independently rated the certainty and resolved disagreements by discussion. The GRADE approach considers five domains that can lower certainty: risk of bias (based on RoB 2 assessments, including the proportion of studies with low-risk, some concerns, or high risk), inconsistency (based on the I2 statistic, prediction intervals, and overlap of confidence intervals), indirectness (based on differences between study populations, interventions, comparators, and outcomes relative to the review question), imprecision (based on optimal information size and whether 95% confidence intervals cross clinically important thresholds, such as an odds ratio of 1.0 for remission or a standardized mean difference of 0.2 for a small effect on depression scores), and publication bias (based on funnel plot asymmetry and Egger's test). Certainty was graded as follows High (Further research is very unlikely to change confidence in the effect estimate); Moderate (Further research is likely to have an important impact on confidence and may change the estimate); Low (Further research is very likely to have an important impact on confidence and is likely to change the estimate); and Very low (Any estimate of effect is very uncertain). Certainty assessments were performed separately for remission (dichotomous outcome) and depression scores (continuous outcome).

Results

Study selection

The PRISMA 2020 flow diagram (Figure 1) summarizes the study selection process. Database searches identified 10,123 records (Cochrane Library: n = 1,845; Embase: n = 2,567; Google Scholar: n = 3,210; MEDLINE via OVID: n = 1,892; PubMed: n = 609). After removing 1,018 duplicate records, 9,105 unique records remained for screening. Discrepancy between stated and applied inclusion criteria for control groups. The innovation section initially stated that studies with active psychological interventions or stretching/relaxation controls were excluded. However, during the review process, excluding active comparators would limit clinical generalizability and remove important studies (e.g., exercise vs. pharmacotherapy). Therefore, the final eligibility criteria (as documented in the Methods) accepted both passive and active non‑exercising controls. This discrepancy is transparently reported here and in the discussion (Limitations). To examine the impact of including active comparators, prespecified subgroup analyses stratifying by control type (passive vs. active; see Table 3) and sensitivity analyses excluding active controls altogether were conducted.

Title and abstract screening

Two reviewers independently screened titles and abstracts of all 9,105 records. A calibration exercise on 100 records (κ = 0.92) confirmed consistent application of eligibility criteria. Of these, 8,997 records were excluded because they clearly did not meet the inclusion criteria;

Not related to depression (n = 3,644); No exercise intervention (n = 3,015); Animal or in vitro studies (n = 556); Reviews, commentaries, editorials, or conference abstracts (n = 1,102); Duplicate publications not identified during the initial EndNote deduplication process but subsequently detected during manual title/abstract screening (n = 153); and Non-English language without available translation (n = 527).

Full-text retrieval

Full-text reports for the remaining 108 records were assessed for eligibility.

Full-text exclusion

After detailed assessment, 68 full-text articles were excluded for the following reasons; No non-exercising control group (e.g., exercise vs. exercise only, or no control) (n = 10); No depression outcome measure (n = 9); No original data (e.g., protocol only, duplicate publication of same cohort) (n = 9); No structured exercise intervention (e.g., physical activity advice only, no prescribed program) (n = 8); Population did not have depression (e.g., healthy volunteers, other psychiatric disorders) (n = 17); Insufficient data for effect size calculation (means/SD not reported and not obtainable from authors) (n = 9); and non-randomized design (observational, uncontrolled) (n = 6).

Final inclusion

Forty randomized controlled trials met all eligibility criteria and were included in the meta-analysis. The characteristics of these 40 studies are presented in Supplementary Table 1. Forty studies were included overall; 37 contributed to remission analysis and 24 to depression severity analysis. A total of 2,667 participants with depression were included across the 40 studies (1,415 in exercise groups, 1,252 in control groups)24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,
58,59,60,61,62,63. The sample size ranged from 16 to 266 participants. These limitations may affect the significance of the findings.

Extraction of continuous outcome data
Across the 40 included studies, the following types of effect estimates were extracted: post-intervention scores only (n = 23, 52.3%), change-from-baseline scores (n = 14, 31.8%), and adjusted effect estimates from ANCOVA or similar models (n = 7, 15.9%). No studies required direct imputation of missing standard deviations using external data (e.g., borrowing SDs from other studies), as all studies provided sufficient raw data—baseline and post-intervention means and SDs—to calculate the required effect estimates. For studies reporting multiple follow-up time points (n = 11), the primary endpoint at the end of the intervention period (median 12 weeks, range 3–16 weeks) was extracted. Long-term follow-up data (≥6 months) were available in 6 studies and are reported separately in a sensitivity analysis. When compared to the control group, exercise demonstrated significantly higher remission (OR, 2.40; 95% CI, 1.97–2.93, p < 0.001) with low heterogeneity (I2 = 49%), this indicates that in a future study conducted in a similar setting, the true OR for remission could be as low as 1.28 (modest benefit) or as high as 4.73 (large benefit), reflecting moderate uncertainty, and lower depression score (SMD, -0.96; 95% CI, -1.27–-0.65, p < 0.001) with high heterogeneity (I2 = 87%) in patients with depression, as shown in Figures 2 and 3. The substantial heterogeneity (I2 = 87%) likely reflects, in part, this scale diversity. This wide interval crosses zero, indicating that in some future settings or populations, exercise might not reduce depression scores or could even be detrimental. This finding underscores the need for individualized prescription of exercise for depression.

Number needed to treat (NNT)

The pooled remission rate in control groups (CER) was 32.4% (430/1326). Using the OR for remission (2.40), the calculated NNT for the primary analysis was 2.1 (95% CI, 1.7–2.8). This means that approximately 2 patients need to receive exercise treatment for each additional patient who achieves remission compared to the control. Using the alternative method based on the primary SMD (converted to Cohen's d = 0.96 in absolute value), the NNT was 2.0, consistent with the OR-based estimate. Subgroup NNT values were as follows; low-risk-of-bias studies (n = 12); NNT = 2.8 (95% CI, 2.1–4.2); MDD-only studies (n = 24); NNT = 1.9 (95% CI, 1.5–2.6); and supervised exercise studies (n = 35); NNT = 1.6 (95% CI, 1.3–2.1)

Comparison with established treatments;

The NNT for exercise in the primary analysis (2.1) was lower (more favorable) than published NNTs for psychotherapy (2.5) and antidepressant medication (4.3). However, these comparisons are indirect and should be interpreted with caution due to differences in study populations, outcome definitions, and time points.

Scale-specific mean differences (BDI and HAM-D)

Using the pooled SMD (–0.96) and the pooled baseline standard deviations from studies using each scale, Beck Depression Inventory (BDI); Based on 8 studies (n = 292) with a pooled baseline SD of 10.3, the estimated mean difference was –6.5 (95% CI, –9.2 to –3.8) BDI points. The observed mean differences in these 8 studies ranged from –4.2 to –8.9, all favoring exercise. Hamilton Depression Rating Scale (HAM-D); Based on 16 studies (n = 634) with a pooled baseline SD of 4.5, the estimated mean difference was –4.7 (95% CI, –6.9 to –2.5) HAM-D points. The observed mean differences in these 16 studies ranged from –3.1 to –6.8, all favoring exercise. Both estimated differences exceed the NICE threshold for clinical significance (≥3 points on BDI or HAM-D), suggesting that exercise produces clinically meaningful reductions in depressive symptoms.

Risk-of-bias assessment results

A summary of the RoB 2 assessments for all 40 included studies is presented in Table 2. Across the five domains, the proportion of studies judged as ‘low-risk of bias,’ ‘some concerns,’ and ‘high risk of bias’ was as follows: risk of bias assessment demonstrated variable methodological quality across the five RoB 2 domains. For Domain 1 (randomization process), 24 studies (60%) were judged as low-risk of bias, 10 studies (25%) had some concerns, and 6 studies (15%) were considered at high risk. In Domain 2 (deviations from intended interventions), 15 studies (37.5%) were rated as low-risk, 15 studies (37.5%) had some concerns, and 10 studies (25%) were classified as high risk. For Domain 3 (missing outcome data), 18 studies (45.0%) were assessed as low-risk, while 12 studies (30%) had some concerns, and 10 studies (25%) were at high risk. Domain 4 (outcome measurement) demonstrated the most favorable ratings, with 27 studies (67.5%) categorized as low-risk, 9 studies (22.5%) as having some concerns, and only 4 studies (10%) as high risk. Finally, for Domain 5 (selection of the reported result), 21 studies (52.5%) were rated as low-risk, 14 studies (35%) had some concerns, and 5 studies (12.5%) were judged to be at high risk of bias. For overall RoB 2 judgments; 12 studies (30%) were rated as ‘low-risk of bias’ (all five domains low-risk), 13 studies (32.5%) as ‘some concerns’ (at least one domain with some concerns but no high-risk domains), and 15 studies (37.5%) as ‘high risk of bias’ (any domain rated high risk or multiple domains with some concerns).

The most common reasons for ‘high risk’ judgments were: inadequate blinding of participants and personnel (Domain 2, 22.7%), high or unbalanced dropout rates without appropriate imputation (Domain 3, 22.7%), and inadequate allocation concealment (Domain 1, 13.6%). Studies published before 2001 (the pre-CONSORT era, n = 5) accounted for 4 of the 15 ‘high risk’ judgments. Inter-rater agreement for overall risk-of-bias judgments was substantial (κ = 0.79; 95% CI, 0.68–0.88). No studies were excluded based on risk-of-bias judgments; instead, sensitivity analyses restricted to studies with ‘low-risk’ or ‘low-risk + some concerns’ are reported below.

Subgroup analyses

Table 3 presents the results of all prespecified subgroup analyses for the remission outcome (OR) and depression score outcome (SMD). Key findings include: Exercise mode as both aerobic (OR, 2.54; 95% CI, 2.01–3.21; n = 24) and resistance (OR, 2.38; 95% CI, 1.72–3.29; n = 8) training, showed significant effects, with no significant difference between modes (p for interaction = 0.68). Mixed aerobic + resistance training showed a smaller effect (OR, 1.92; 95% CI, 1.31–2.81; n = 6), though the interaction was not significant (p = 0.21). Supervision as supervised exercise (OR, 2.71; 95% CI, 2.18–3.36; n = 31) produced larger effects than unsupervised exercise (OR, 1.88; 95% CI, 1.34–2.64; n = 9), and the difference was statistically significant (p for interaction = 0.04). Control group type as passive controls (waitlist/usual care) yielded larger effects (OR, 2.89; 95% CI, 2.31–3.61; n = 24) than active controls (active non-exercise controls were allowed) (attention control/pharmacotherapy/psychotherapy; OR, 1.79; 95% CI, 1.34–2.38; n = 16; p for interaction = 0.02). Depression diagnosis, as studies requiring a formal MDD diagnosis (OR, 2.32; 95% CI, 1.82–2.95; n = 20) showed similar effects to studies using screening tool cutoffs (OR, 2.61; 95% CI, 1.99–3.42; n = 20; p for interaction = 0.41). Risk of bias as low-risk-of-bias studies (OR, 2.08; 95% CI, 1.52–2.85; n = 12) showed smaller effects than studies with some concerns (OR, 2.68; 95% CI, 1.94–3.70; n = 13) or high risk (OR, 2.54; 95% CI, 1.81–3.56; n = 15), but the differences were not statistically significant (p for interaction = 0.34). Sample size as larger studies (n ≥ 100) showed smaller effects (OR, 1.96; 95% CI, 1.45–2.65; n = 8) than medium (OR, 2.61; 95% CI, 1.99–3.42; n = 16) or small studies (OR, 2.58; 95% CI, 1.86–3.58; n = 16), suggesting potential small-study effects, though the interaction was not significant (p = 0.28). Egger's test did not detect publication bias (p = 0.87), but this pattern is consistent with some degree of bias or heterogeneity.

For depression scores (SMD), findings were largely consistent with those for remission, although wide confidence intervals and high heterogeneity limited interpretability for several subgroups. Detailed results for all subgroup analyses are provided in Table 3. Although aerobic exercise was the most frequently studied modality (28 trials), resistance training and mixed protocols also showed significant effects, with overlapping confidence intervals suggesting no clear superiority of one modality over another.

Certainty of evidence (GRADE)

Using the GRADE framework, the certainty of evidence for depression remission was rated as moderate (Table 3). Certainty was downgraded by one level for serious risk of bias (37.5% of studies with overall ‘high risk’ per RoB 2). Inconsistency was not serious (I2 = 48%; prediction interval did not cross 1.0), indirectness was not present, imprecision was not serious (optimal information size met), and publication bias was not detected. This indicates moderate confidence that exercise increases depression remission compared to control conditions; however, further high-quality research may change this estimate. For depression severity scores (SMD), the certainty of evidence was rated as low (Table 4). Certainty was downgraded by one level for serious risk of bias (same reason as above) and by an additional level for very serious inconsistency (I2 = 87%; 95% prediction interval −1.42 to 0.20, which crosses the null effect). The wide prediction interval indicates that in some future settings, exercise might not reduce depression scores or could even be detrimental. Low certainty means that confidence in the effect estimate is limited, and the true effect may be substantially different from the point estimate.

Sensitivity analysis by extraction type

To assess whether the choice of extraction method influenced the pooled effect size, studies were stratified by the type of effect estimate extracted. There was no statistically significant difference between these subgroups (p for interaction = 0.54), suggesting that the extraction method did not systematically bias the results. However, all subgroups exhibited substantial heterogeneity, with the lowest heterogeneity observed in the adjusted-effects subgroup (I2 = 72%).

The use of stratified models to inspect the effects of certain components was not possible owing to the lack of data, such as gender, ethnicity, and age, on comparison outcomes. Visual inspection of the funnel plot for remission (Figure 4) showed approximate symmetry. The Egger regression test for remission yielded a p-value of 0.87, indicating no statistically significant evidence of small-study effects or publication bias for this outcome. For depression scores (Figure 5), the funnel plot appeared roughly symmetrical, with an Egger test p-value of 0.68, similarly suggesting no evidence of publication bias. However, given the high heterogeneity for depression scores (I2 = 87%), funnel plot asymmetry due to heterogeneity rather than bias cannot be ruled out. The RoB 2 assessment revealed that 15 of 40 studies (37.5%) were judged as having an overall ‘high risk of bias,’ primarily due to inadequate blinding of participants and personnel (Domain 2), incomplete outcome data (Domain 3), and inadequate allocation concealment (Domain 1). An additional 13 studies (32.5%) had ‘some concerns.’ Only 12 studies (30%) were rated as overall ‘low-risk of bias.’ Selective reporting (Domain 5) was judged as low-risk for the majority of studies (56.8%), suggesting that reporting bias was not a major concern.

Handling of multi-arm trials

Of the 40 included studies, 9 (20.5%) had multi-arm designs requiring special handling. For studies with multiple exercise arms compared to a single control group (n = 5), exercise arms were combined to avoid double-counting control participants. For studies with both exercise and multiple non-exercise comparator arms (n = 4), the most conservative non-exercising control was selected (prioritizing waitlist or usual care over pharmacotherapy or psychotherapy). A sensitivity analysis excluding all multi-arm studies yielded similar results to the primary analysis (remission: OR, 2.38; 95% CI, 1.94–2.91 vs. primary OR 2.46; depression SMD: −0.58; 95% CI, −0.87 to −0.29 vs. primary SMD −0.96), indicating that the handling of multi-arm trials did not substantially bias the findings.

DATA AVAILABILITY:

The datasets analyzed during the current study are available from

https;//zenodo.org/records/19980462. The new PRISMA checklist is attached as Supplementary File 1.

FIGURES AND TABLES LEGENDS:

Systematic review process flowchart detailing study selection for meta-analysis.
Figure 1: PRISMA flow diagram of the study selection process. The diagram shows the number of records identified, screened, assessed for eligibility, and included in the meta-analysis. Boxes represent sequential stages of the review process: identification, screening, eligibility assessment, and inclusion. Arrows indicate the flow of records through each stage. Reasons for exclusion and corresponding numbers are provided in the exclusion boxes. Abbreviations: n = number of records or studies; PRISMA = preferred reporting items for systematic reviews and meta-analyses. Please click here to view a larger version of this figure.

Forest plot diagram, odds ratio analysis comparing exercise interventions vs. control in studies.
Figure 2: Forest plot of the effect of exercise versus control on depression remission (odds ratio). Each horizontal line represents an individual study (n = 37). The square at the center of each line represents the point estimate (odds ratio, OR), with the size of the square proportional to the study's weight in the pooled analysis (inverse variance method) under a random‑effects model (DerSimonian and Laird). The horizontal line represents the 95% confidence interval (CI). The vertical dashed line at OR = 1.0 represents the line of no effect. The diamond at the bottom represents the pooled effect estimate; the center of the diamond is the pooled OR (2.40), and the width represents the 95% CI (1.97 to 2.93). Heterogeneity statistics (I2 = 49%) and the p‑value for the overall effect (p < 0.001) are shown below the diamond. No fixed‑effect model was used for this primary analysis. Abbreviations: OR = odds ratio; CI = confidence interval; Random‑effects model (DerSimonian and Laird). Please click here to view a larger version of this figure.

Meta-analysis forest plot of exercise intervention effect, showing standard mean difference results.
Figure 3: Forest plot of the effect of exercise versus control on depression severity (standardized mean difference, SMD). Each horizontal line represents an individual study (n = 24). The square at the center of each line represents the point estimate (SMD, Hedges' g), with the size of the square proportional to the study's weight in the pooled analysis (inverse variance method). The horizontal line represents the 95% confidence interval (CI). Negative SMD values favor exercise (reduced depression severity). The vertical solid line at SMD = 0 represents the line of no effect. The diamond at the bottom represents the pooled effect estimate using a random-effects model (DerSimonian and Laird method); the center of the diamond is the pooled SMD (−0.96), and the width represents the 95% CI (−1.27 to −0.65). Heterogeneity statistics (I2 = 87%, tau2 = 0.46) and the p-value for the overall effect (p < 0.001) are shown below the diamond. Abbreviations: SMD = standardized mean difference (Hedges' g); CI = confidence interval; IV = inverse variance; Random = random-effects model. Please click here to view a larger version of this figure.

Funnel plot diagram showing SE(log[OR]) vs OR, used for meta-analysis publication bias detection.
Figure 4: Funnel plot for remission (odds ratio). Each point represents an individual study (n = 37). The vertical line represents the pooled effect estimate (OR = 2.40). The diagonal lines represent pseudo-95% confidence intervals. The approximate symmetry of the plot and Egger's test (p = 0.87) suggests no evidence of publication bias or small-study effects. Please click here to view a larger version of this figure.

Funnel plot showing standard error vs. standard mean difference in statistical data analysis.
Figure 5: Funnel plot assessing potential publication bias for the depression severity outcome (standardized mean difference, SMD). Each circle represents an individual study comparison (n = 24). The vertical line indicates the pooled effect estimate (SMD = −0.96). Studies are plotted by effect size (x-axis) and standard error (SE; y-axis). Larger studies with smaller standard errors appear toward the top of the plot, whereas smaller studies appear toward the bottom. Visual inspection suggests approximate symmetry, although dispersion is evident because of substantial between-study heterogeneity (I2 = 87%). Egger's regression test was not statistically significant (intercept = −0.25, 95% CI −1.08 to 0.58; p = 0.68), indicating no statistical evidence of small-study effects. Abbreviations: SMD = standardized mean difference; SE = standard error; CI = confidence interval. Please click here to view a larger version of this figure.

Database (platform)Search dateComplete search syntaxLimits/filters
PubMedNov-25("exercise"[MeSH Terms] OR "exercise"[tiab] OR "physical activity"[tiab] OR "aerobic"[tiab] OR "resistance training"[tiab] OR "strength training"[tiab] OR "yoga"[tiab] OR "walking"[tiab]) AND ("depression"[MeSH Terms] OR "depressive disorder"[MeSH Terms] OR "depression"[tiab] OR "depressive symptoms"[tiab] OR "major depressive disorder"[tiab] OR "dysthymia"[tiab]) AND ("remission"[MeSH Terms] OR "remission"[tiab] OR "antidepressant agents"[MeSH Terms] OR "medication"[tiab] OR "pharmacotherapy"[tiab] OR "treatment outcome"[MeSH Terms])Humans; English language; no date limits
Embase (OVID)Nov-25(exp exercise/ or exp physical activity/ or (exercise$ or physical activity$ or aerobic$ or resistance training or strength training or yoga or walking).ti,ab.) AND (exp depression/ or exp major depression/ or (depression$ or depressive symptoms or major depressive disorder or dysthymia).ti,ab.) AND (exp remission/ or exp drug therapy/ or (remission$ or medication$ or pharmacotherapy$ or treatment outcome$).ti,ab.)Humans; English language; articles in press; no date limits
Cochrane Central Register of Controlled Trials (CENTRAL)Nov-25#1 MeSH descriptor: [Exercise] explode all trees OR MeSH descriptor: [Physical Activity] explode all trees OR (exercise* OR "physical activity" OR aerobic OR "resistance training" OR "strength training" OR yoga OR walking):ti,ab,kwTrials only (CENTRAL default); no date limits
#2 MeSH descriptor: [Depression] explode all trees OR MeSH descriptor: [Depressive Disorder] explode all trees OR (depression* OR "depressive symptoms" OR "major depressive disorder" OR dysthymia):ti,ab,kw
#3 MeSH descriptor: [Remission] explode all trees OR MeSH descriptor: [Antidepressive Agents] explode all trees OR (remission* OR medication* OR pharmacotherapy* OR "treatment outcome"):ti,ab,kw
#4 #1 AND #2 AND #3
MEDLINE (OVID)Nov-25(exp exercise/ or exp physical activity/ or (exercise$ or physical activity$ or aerobic$ or resistance training or strength training or yoga or walking).ti,ab.) AND (exp depression/ or exp depressive disorder/ or (depression$ or depressive symptoms or major depressive disorder or dysthymia).ti,ab.) AND (exp remission/ or exp antidepressive agents/ or (remission$ or medication$ or pharmacotherapy$ or treatment outcome$).ti,ab.)Humans; English language; randomized controlled trials; no date limits
Google ScholarNov-25"exercise" AND "depression" AND ("remission" OR "medication" OR "treatment")First 500 records sorted by relevance; no date limits; English only
Additional sourcesNov-25Reference lists of included studies and relevant reviews (snowball searching); hand-searching of key journals (Mental Health and Physical Activity, Journal of Affective Disorders, Depression and Anxiety)Not applicable

Table 1: Database search strategies. Complete search syntax used for each database, including controlled vocabulary (MeSH, Emtree), keywords, field tags, Boolean operators, and truncation. Searches were conducted in November 2025. Limits applied: human studies, English language, no date restrictions (except where noted). Legend: #1, #2, #3 = sequential search set numbers representing successive search steps. Truncation symbols were used to capture variant word endings (e.g., exercise retrieves exercise, exercises, and exercising): $ in Ovid and * in PubMed and Cochrane. For Ovid searches, the database searched was MEDLINE (Ovid platform). Due to Google Scholar's limitations on complex Boolean searching, a simplified search string was used; results were restricted to the first 500 records sorted by relevance. The Google Scholar search was conducted on 15 November 2025 at 10:00 AM GMT, with personalization disabled. Abbreviations: MeSH = Medical Subject Headings (PubMed); /exp = explode (Emtree terms in Embase); ti, ab = title and abstract; kw = keywords (Cochrane); kw = author keywords (OVID); All Fields = unqualified search across all indexed fields. Please click here to download this Table.

StudyD1: RandomizationD2: DeviationsD3: Missing dataD4: Outcome measurementD5: Reported resultOverall
Nimmo, 198624Some concernsHigh riskHigh riskSome concernsSome concernsHigh risk
Doyne, 198725Some concernsHigh riskSome concernsSome concernsSome concernsHigh risk
McNeil, 199126Some concernsHigh riskHigh riskSome concernsSome concernsHigh risk
Veale, 199227High riskHigh riskSome concernsHigh riskHigh riskHigh risk
Singh, 199728High riskHigh riskHigh riskHigh riskSome concernsHigh risk
Mather, 200229High riskHigh riskSome concernsHigh riskSome concernsHigh risk
Nabkasorn, 200530Low riskLow riskHigh riskLow riskHigh riskHigh risk
Singh, 200531Low riskSome concernsLow riskLow riskSome concernsSome concerns
Sims, 200632Some concernsSome concernsHigh riskSome concernsSome concernsHigh risk
Brenes, 200733Some concernsHigh riskSome concernsSome concernsSome concernsHigh risk
Blumenthal, 200734Low riskLow riskLow riskLow riskLow riskLow risk
Pilu, 200735High riskHigh riskHigh riskSome concernsHigh riskHigh risk
Vieira, 200736High riskHigh riskHigh riskSome concernsHigh riskHigh risk
Williams, 200837Some concernsSome concernsHigh riskSome concernsSome concernsHigh risk
Sims, 200938Low riskSome concernsSome concernsLow riskSome concernsSome concerns
Shahidi, 201139Some concernsSome concernsLow riskLow riskSome concernsSome concerns
Mota-Pereira, 201140Low riskSome concernsSome concernsLow riskLow riskSome concerns
Prakhinkit, 201341Low riskLow riskLow riskLow riskLow riskLow risk
Legrand, 201442Some concernsSome concernsLow riskLow riskLow riskSome concerns
Pfaff, 201443Low riskLow riskLow riskLow riskLow riskLow risk
Danielsson, 201444Low riskSome concernsSome concernsLow riskLow riskSome concerns
Oertel-Knöchel, 201445High riskHigh riskHigh riskSome concernsHigh riskHigh risk
Hallgren, 201546Low riskLow riskLow riskLow riskLow riskLow risk
Carneiro, 201547Low riskSome concernsSome concernsLow riskLow riskSome concerns
Doose, 201548Low riskSome concernsSome concernsLow riskLow riskSome concerns
Schuch, 201549Low riskLow riskLow riskLow riskLow riskLow risk
Gao, 201650Low riskSome concernsSome concernsLow riskSome concernsSome concerns
Schneider, 201651Some concernsSome concernsHigh riskLow riskSome concernsHigh risk
Lok, 201752Low riskLow riskLow riskHigh riskLow riskHigh risk
Cheung, 201853Low riskLow riskSome concernsLow riskLow riskSome concerns
Roy, 201854Low riskSome concernsSome concernsLow riskLow riskSome concerns
Abdelbasset, 201955Low riskLow riskLow riskLow riskLow riskLow risk
Makizako, 202056Low riskLow riskLow riskLow riskLow riskLow risk
La Rocque, 202157Low riskLow riskLow riskLow riskLow riskLow risk
Chau, 202258Low riskLow riskLow riskLow riskLow riskLow risk
Nicholas, 202459Some concernsSome concernsLow riskLow riskSome concernsSome concerns
Woolf, 202460Low riskLow riskLow riskLow riskLow riskLow risk
James Vibin, 202461Low riskSome concernsLow riskLow riskLow riskSome concerns
Chang, 202562Low riskLow riskLow riskLow riskLow riskLow risk
Wei, 202563Low riskLow riskLow riskLow riskLow riskLow risk

Table 2: Risk-of-bias assessment for included studies using RoB 2. This table presents the risk‑of‑bias judgments for each of the 40 included randomized controlled trials, assessed using the Cochrane RoB 2 tool (version August 2019). Five domains are evaluated for each study: D1 (randomization process), D2 (deviations from intended interventions), D3 (missing outcome data), D4 (outcome measurement), and D5 (selection of the reported result). For each domain, the judgment is classified as “low-risk of bias”, “some concerns”, or “high risk of bias”. The overall risk‑of‑bias judgment is derived according to RoB 2 algorithms: “low-risk” if all domains are low-risk, “some concerns” if at least one domain has some concerns but no high‑risk domains, and “high risk” if any domain is rated high risk or multiple domains have some concerns. Inter‑rater agreement for overall judgments was substantial (κ = 0.79). Studies were not excluded based on risk of bias; instead, sensitivity analyses restricted to low‑risk and low‑risk‑plus‑some‑concerns studies were performed. Abbreviations: D1 = Randomization process; D2 = Deviations from intended interventions; D3 = Missing outcome data; D4 = Outcome measurement; D5 = Selection of reported result. Please click here to download this Table.

Subgroup variableSubgroup levelNo. of studiesEffect estimate (OR or MD) [95% CI]I² (%)p for interactionPooled SMD (95% CI)
Exercise modeAerobic24OR 2.54 [2.01–3.21]44%0.68−0.92 (−1.28 to −0.56)
Resistance8OR 2.38 [1.72–3.29]41%−1.04 (−1.45 to −0.63)
Mixed6OR 1.92 [1.31–2.81]52%−0.78 (−1.18 to −0.38)
Other (yoga/dance)2OR 2.15 [1.23–3.76]0%−0.85 (−1.32 to −0.38)
SupervisionSupervised31OR 2.71 [2.18–3.36]45%0.04−1.12 (−1.48 to −0.76)
Unsupervised9OR 1.88 [1.34–2.64]39%−0.65 (−1.02 to −0.28)
Control typePassive (waitlist/usual care/no intervention)24OR 2.89 [2.31–3.61]42%0.02−1.18 (−1.56 to −0.80)
Active (attention/pharmacotherapy/psychotherapy)16OR 1.79 [1.34–2.38]51%−0.68 (−1.05 to −0.31)
Depression diagnosisFormal MDD diagnosis20OR 2.32 [1.82–2.95]49%0.41−0.98 (−1.35 to −0.61)
Screening tool cutoff20OR 2.61 [1.99–3.42]46%−0.92 (−1.30 to −0.54)
Risk of bias (RoB 2)Low risk12OR 2.08 [1.52–2.85]38%0.34−0.48 (−0.82 to −0.14)
Some concerns13OR 2.68 [1.94–3.70]44%−1.06 (−1.45 to −0.67)
High risk15OR 2.54 [1.81–3.56]51%−1.14 (−1.52 to −0.76)
Sample size (total N)Small (<50)16OR 2.58 [1.86–3.58]48%0.28−1.08 (−1.48 to −0.68)
Medium (50–99)16OR 2.61 [1.99–3.42]46%−0.96 (−1.34 to −0.58)
Large (≥100)8OR 1.96 [1.45–2.65]35%−0.72 (−1.10 to −0.34)
Age groupAdults (18–59 years)20OR 2.44 [1.94–3.07]47%0.92−0.94 (−1.32 to −0.56)
Older adults (≥60 years)20OR 2.48 [1.88–3.27]49%−0.98 (−1.36 to −0.60)
Exercise durationShort (≤8 weeks)16OR 2.61 [2.01–3.39]43%0.58−1.02 (−1.40 to −0.64)
Medium (9–12 weeks)18OR 2.38 [1.84–3.08]49%−0.94 (−1.32 to −0.56)
Long (≥13 weeks)6OR 2.28 [1.51–3.44]52%−0.82 (−1.22 to −0.42)
Publication yearPre-201014OR 2.32 [1.76–3.06]51%0.67−0.88 (−1.26 to −0.50)
2010–201917OR 2.51 [1.93–3.26]46%−0.98 (−1.36 to −0.60)
2020–20259OR 2.54 [1.88–3.43]44%−1.02 (−1.40 to −0.64)

Table 3: Subgroup analysis results. This table displays the results of prespecified subgroup analyses examining potential moderators of exercise efficacy on depression remission (odds ratio). Subgroup variables include exercise characteristics (mode, supervision, duration), population characteristics (depression diagnosis type, age group, baseline severity), and study design features (control group type, risk of bias, sample size, publication year). For each subgroup level, the table reports the number of studies, the pooled odds ratio with 95% confidence interval, the I2 statistic for heterogeneity, and the p‑value for interaction between subgroups. A p‑value for interaction of less than 0.10 was considered statistically significant, acknowledging the low power of interaction tests in meta‑analyses. Statistically significant subgroup differences (p < 0.10) are bolded in the table. Control group type was categorized as passive (waitlist, usual care, no intervention) or active (attention control, pharmacotherapy, psychotherapy). Risk of bias was assessed using the Cochrane RoB 2 tool. All subgroup analyses used random‑effects models (DerSimonian and Laird method). Bolded p for interaction values indicates statistically significant subgroup differences (p < 0.10). Please click here to download this Table.

OutcomeNo. of participants (Studies)Effect estimate (95% CI)Certainty assessmentGRADE certainty
Depression remission2667 (40 RCTs)OR 2.40 (1.97–2.93)Risk of bias: Serious (−1) – 37.5% of studies high riskModerate Static equilibrium, ΣFx=0, ΣFy=0, diagram, forces balance, vector analysis, engineering principles.Static equilibrium, ΣFx=0, ΣFy=0, diagram, forces balance, vector analysis, engineering principles.Static equilibrium, ΣFx=0, ΣFy=0, diagram, forces balance, vector analysis, engineering principles.
Inconsistency: Not serious – I² = 48%, prediction interval 1.28–4.73 (does not cross 1.0)
Indirectness: Not serious – direct comparison
Imprecision: Not serious – OIS met, CI excludes no effect
Publication bias: Not detected (Egger p = 0.87)
Depression severity (SMD)2667 (40 RCTs)SMD −0.96 (−1.27−0.65)Risk of bias: Serious (−1) – 37.5% of studies high riskLow Static equilibrium, ΣFx=0, ΣFy=0, diagram, forces balance, vector analysis, engineering principles.Static equilibrium, ΣFx=0, ΣFy=0, diagram, forces balance, vector analysis, engineering principles.◯◯
Inconsistency: Serious (−1) – I² = 87%, prediction interval −1.42 to 0.20 (crosses zero)
Indirectness: Not serious – direct comparison
Imprecision: Not serious – OIS met
Publication bias: Not detected (Egger p = 0.68)

Table 4: Summary of findings (GRADE) – Exercise compared to control for depression symptoms. This table presents the Grading of Recommendations Assessment, Development and Evaluation (GRADE) certainty of evidence for the two primary outcomes: depression remission (dichotomous) and depression severity (continuous, reported as standardized mean difference). For each outcome, the table shows the number of participants and studies, the pooled effect estimate with 95% confidence interval, and a structured certainty assessment across five domains: risk of bias (based on RoB 2 assessments), inconsistency (based on I2 statistic and prediction intervals), indirectness (based on population, intervention, comparator, and outcome differences), imprecision (based on optimal information size and confidence interval crossing of clinically important thresholds), and publication bias (based on funnel plot asymmetry and Egger's test). Certainty is graded as high, moderate, low, or very low. For remission, certainty was downgraded by one level for serious risk of bias, resulting in moderate certainty. For depression severity scores, certainty was downgraded by one level for serious risk of bias and an additional level for very serious inconsistency (I2 = 87% and prediction interval crossing zero), resulting in low certainty. Explanation of certainty ratings: Moderate certainty (remission): authors are moderately confident that the true effect is likely close to the estimated effect, although further research may change this estimate. Low certainty (depression scores): authors have limited confidence in the estimated effect; the true effect may differ substantially from the estimate, as reflected by the wide prediction interval crossing the line of no effect and the very high statistical heterogeneity (inconsistency). Abbreviations: OR = odds ratio; SMD = standardized mean difference; CI = confidence interval; OIS = optimal information size; GRADE = Grading of Recommendations Assessment, Development and Evaluation. Please click here to download this Table.

Supplementary Table 1: Characteristics of the 40 included studies in the meta-analysis. Studies are ordered chronologically by publication year. All values represent the number of participants (n) unless otherwise specified. Total sample sizes include all randomized participants (intention-to-treat populations where reported). For studies with multiple exercise arms, exercise group counts represent the sum of participants across exercise arms, and control group counts represent the original control arm (not divided) to avoid double-counting (see Methods for handling of multi-arm trials). Legend: Country names are reported as presented in the original studies. Exercise (n) = number of participants randomized to the exercise intervention group(s); for studies with multiple exercise arms, n represents the combined sample size across all exercise arms. Control (n) = number of participants randomized to the control (non-exercising) group; for studies with multiple control arms, only the non-exercising control group was included, using the following prioritization hierarchy: waitlist > usual care > attention control > pharmacotherapy > psychotherapy. Sample sizes reflect the number of participants randomized at baseline and do not account for attrition during follow-up. For studies published before 2001 (pre-CONSORT era; n = 5: Nimmo 1986, Doyne 1987, McNeil 1991, Veale 1992, and Singh 1997), sample sizes are reported as stated in the original publications. No studies required imputation of missing sample size data. Control type classification: Passive = waitlist, usual care, or no intervention; Active = attention control, pharmacotherapy, or psychotherapy. Studies with multiple control arms were handled according to the hierarchical selection approach described in the Methods, with passive controls prioritized over active controls. Abbreviations: N = total sample size (exercise + control participants); UK = United Kingdom; USA = United States of America; N/A = not applicable.Please click here to download this file.

Supplementary File 1: NEW PRISMA checklist.Please click here to download this file.

Discussion

This meta‑analysis of 40 randomized controlled trials (2,667 participants with depression)24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,
58,59,60,61,62,63, found that exercise significantly increased remission rates (OR = 2.40; 95% CI, 1.97–2.93; I2 = 49%) and reduced depression severity scores (SMD = −0.96; 95% CI, −1.27 to −0.65; I2 = 87%) compared to non‑exercising controls. The number needed to treat for one additional remission was 2.1, and the pooled SMD translated into clinically meaningful reductions (≈6.5 Beck Depression Inventory points, ≈4.7 Hamilton Depression Rating Scale points). These findings provide preliminary support for exercise as an effective complementary intervention for depression, particularly for achieving remission. However, the very high heterogeneity for continuous outcomes (I2 = 87%) and the fact that only 12 of 40 studies (30%) were rated as low-risk of bias warrant cautious interpretation. Substantial heterogeneity for depression scores arises from multiple sources: variations in exercise protocols, intervention duration (3–16 weeks), supervision status, control group types (waitlist, usual care, pharmacotherapy, psychotherapy), baseline depression severity, and outcome scales15,16. The included studies employed diverse modalities—aerobic, resistance, mixed, yoga, tai chi, and dance—which may differ in antidepressant mechanisms, acceptability, and adherence16.

Subgroup analyses showed no clear superiority of any single modality, but supervised exercise produced larger effects than unsupervised (p = 0.04), and passive controls (waitlist/usual care) yielded larger effects than active comparators (p = 0.02). These findings indicate that both the nature of the intervention and the choice of comparator influence observed effect sizes. Light‑intensity and low‑burden exercise options warrant further consideration, as many patients with depression have low baseline fitness, fatigue, and low motivation64,65. Future studies should test whether such interventions produce clinically meaningful effects in populations unable or unwilling to engage in moderate‑to‑vigorous intensity programs. These results align with Schuch et al. (2016; SMD = −0.62)10,11,12 but extend the evidence base by including 23 additional trials published since 2019. Unlike Krogh et al. (2017)12, who found no effect when restricting to low‑risk‑of‑bias trials, the present sensitivity analysis, limited to low-risk studies, still showed a significant but attenuated effect (SMD = −0.48). This discrepancy may reflect the use of the updated RoB 2 tool and the inclusion of more recent high‑quality trials. The number needed to treat (NNT = 2.1) compares favorably with published NNTs for psychotherapy (2.5)66 and antidepressant medication (4.3)67, although these are indirect comparisons. NNT for medication is derived from placebo‑controlled trials with strict blinding, whereas NNT for exercise comes from open‑label trials; the apparent superiority of exercise may reflect expectancy effects or heterogeneity in control conditions rather than true biological superiority. Thus, the NNT comparison is intended for hypothesis generation, not to claim that exercise is better than medication. Only six studies reported outcomes at ≥6 months, and these showed attenuated effects (SMD = −0.31 vs. −0.96 post‑intervention), indicating that benefits may diminish without ongoing intervention. Clinicians should consider exercise as an adjunct to established treatments, particularly for motivated68, physically able patients with access to supervised programs69.

Several important limitations are present. First, eligibility criteria were not fully pre‑specified, and a protocol (e.g., PROSPERO) was not registered, increasing the risk of reporting bias. Second, the very high heterogeneity for depression severity (I2 = 87%) persisted despite using SMD; the 95% prediction interval crossed zero, meaning exercise may not be beneficial in some future settings or populations. Third, most studies had short follow‑up (≤16 weeks), small sample sizes (only eight studies with ≥100 participants), and were conducted in high‑income countries, limiting generalizability to low‑resource settings. Fourth, placebo effects cannot be excluded due to inherent difficulties in blinding participants to exercise versus control. Fifth, only 15.9% of studies provided adjusted effect estimates (ANCOVA); more than half contributed only post‑intervention scores without baseline adjustment. Sixth, for studies reporting change scores, a correlation coefficient of r = 0.5 was assumed when calculating standard deviations, which may not reflect true within‑study correlations. Finally, adherence barriers related to depression symptoms—low motivation, fatigue, anhedonia—are rarely captured in efficacy trials but are highly relevant for real‑world translation70.

A notable inconsistency in the inclusion criteria should be acknowledged. Although the Innovation section stated that studies with active psychological interventions or stretching/relaxation controls were excluded, the eligibility criteria in the Methods accepted pharmacotherapy and psychotherapy as comparators, and Supplementary Table 1 includes several such studies. This discrepancy arose from an evolving definition of “non‑exercising control” during the review process. Sensitivity analyses stratifying by control type (Table 3) showed significantly larger effects for passive controls (OR = 2.89) than for active controls (OR = 1.79; p = 0.02), indicating that inclusion of active comparators attenuates the overall estimate. Readers should interpret the pooled effect sizes as reflecting a mix of passive and active comparators. Future research should stratify analyses by control condition type or restrict comparisons to non‑active controls only. Additionally, digital and remote exercise interventions (online programs, telehealth coaching, wearable monitors) represent an emerging and potentially scalable approach to managing depression71,72, but none of the trials in this meta‑analysis evaluated these formats. Future RCTs should directly compare digitally delivered exercise to in‑person exercise and to non‑exercising controls71,72.

Exercise was associated with significantly higher remission rates and lower depression scores in patients with depression, but the evidence is limited by high heterogeneity, short follow‑up, and methodological concerns. The GRADE certainty was moderate for remission but low for depression severity scores. Therefore, exercise should currently be recommended as an effective complementary intervention rather than a standalone first‑line treatment. High‑quality, large‑scale randomized controlled trials with longer follow-up, standardized outcome measures, diverse populations, and prespecified subgroup analyses by exercise modality are urgently needed. Future trials should also examine which patients benefit most from which exercise regimen, test light‑intensity and remote delivery formats, and report adverse events and adherence barriers. In the meantime, clinicians should consider patient preferences, physical capabilities, and access to resources when prescribing exercise as part of a comprehensive depression treatment plan73,74.

In conclusion, exercise was associated with significantly higher remission rates (low heterogeneity) and a significantly lower depression score (very high heterogeneity) in patients with depression. Given the substantial heterogeneity for continuous depression outcomes and the methodological limitations of many included studies, these findings provide preliminary rather than definitive support for exercise in depression management. High-quality, large-scale randomized controlled trials with standardized outcome measures are needed before exercise can be recommended as a stand-alone treatment. Clinicians should consider exercise as a potential complementary intervention while awaiting more robust evidence. An additional limitation concerns the pooling of depression scores. Included studies used at least six different depression rating scales (HAM-D, BDI, CES-D, PHQ-9, GDS, and others), each with different ranges and clinical interpretations. Future meta-analyses should consider pooling only studies that use the same scale or performing sensitivity analyses stratified by scale type. The very high heterogeneity (I2 = 87%) underscores the impact of measurement diversity on pooled estimates. Adherence barriers related to depression symptoms deserve specific attention. Depression-related symptoms such as low motivation, fatigue, anhedonia, and reduced perceived capacity for effort may substantially affect exercise initiation and adherence in real-world settings. These factors are rarely captured in efficacy trials but are highly relevant for translating findings into clinical and community practice. Designing feasible, low-burden exercise programs that account for these motivational and effort-perception barriers is essential. However, none of the trials in this meta‑analysis evaluated these formats, and the current evidence base for remote delivery is limited by small sample sizes, short follow‑up, and lack of active comparators. Home-based, wearable-supported, or remotely coached exercise models may improve accessibility and scalability, particularly for patients who cannot access supervised facility-based programs. Future RCTs should directly compare digitally delivered exercise to in‑person exercise and to non‑exercising controls, while addressing barriers such as the digital divide and adherence. Until then, remote exercise should be viewed as a complementary option, not a replacement for supervised programs.

Disclosures

The authors have no competing interests to declare.

Acknowledgements

Not applicable.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
Assumed correlation coefficient for change-score SDImputation of Missing DataCochrane Handbook defaultr = 0.5 (sensitivity analysis performed)
ClinicalTrials.govClinical Trial RegistriesU.S. National Library of MedicineDatabase as of November 10, 2025 (no NCT-specific filter)
Cochrane Central Register of Controlled Trials (CENTRAL)DatabasesCochrane LibraryIssue 11 of 12, November 2025
Cochrane Handbook formulasFormulas & CalculationsCochrane CollaborationHandbook version 6.4 (updated September 2024)
Egger’s linear regression testFunnel Plot AsymmetryRevMan 5.3 / custom macroN/A
Embase (OVID platform)DatabasesElsevier / OVID TechnologiesEmbase Classic+Embase 1947 to 2025 Week 47
EndNoteReference ManagementClarivate Analytics (London, UK)Version 20 (build 20.1.0.16478)
Google ScholarDatabasesGoogle LLCSearch conducted November 15–16, 2025 (first 500 records by relevance)
GRADE (Grading of Recommendations Assessment, Development and Evaluation)Certainty of EvidenceGRADE Working GroupGRADEpro GDT (online version, accessed November 2025)
MEDLINE (OVID platform)DatabasesOVID TechnologiesOvid MEDLINE® 1946 to November 14, 2025
NNT from SMD (Furukawa & Leucht method)Effect Size ConversionBased on area under normal curveN/A
Number needed to treat (NNT) from OREffect Size ConversionFormula: NNT = 1 / [CER × (1 – OR-¹)]N/A
PubMedDatabasesU.S. National Library of MedicineNLM ID: 101775772 (database version 2025)
Review Manager (RevMan)Statistical & Meta-Analysis SoftwareCochrane CollaborationVersion 5.3 (Copenhagen: The Nordic Cochrane Centre)
RoB 2 toolRisk of Bias AssessmentCochrane CollaborationVersion: August 2019 (latest update)
Standardized mean difference (Hedges’ g)Effect Size ConversionCustom calculations in RevMan 5.3N/A
WHO International Clinical Trials Registry Platform (ICTRP)DatabasesWorld Health OrganizationSearch portal version 1.6 (accessed November 2025)

References

  1. James SL, et al. Global, regional, and national incidence, prevalence, and years lived with disability for 354 diseases and injuries for 195 countries and territories, 1990&2013;2017: a systematic analysis for the Global Burden of Disease Study 2017. The Lancet. 2018;392(10159):1789-858.
  2. Organization WH. Depression and other common mental disorders: global health estimates. World Health Organization; 2017.
  3. Malhi GS, et al. The 2020 Royal Australian and New Zealand College of Psychiatrists clinical practice guidelines for mood disorders. Australian & New Zealand Journal of Psychiatry. 2021;55(1):7-117.
  4. Cuijpers P, et al. Psychological Treatment of Depression Compared With Pharmacotherapy and Combined Treatment in Primary Care: A Network Meta-Analysis. The Annals of Family Medicine. 2021;19(3):262-70.
  5. Wise J. NICE guidance on depression: 35 health organizations demand “full and proper” revision. BMJ. 2019;365:l2356.
  6. Thornicroft G, et al. Undertreatment of people with major depressive disorder in 21 countries. British Journal of Psychiatry. 2017;210(2):119-24.
  7. Kessler RC. The costs of depression. Psychiatr Clin North Am. 2012;35(1):1-14.
  8. Activity MP, Team W. WHO guidelines on physical activity and sedentary behavior. World Health Organization. 2020.
  9. McAllister-Williams RH. NICE guidelines for the management of depression. British Journal of Hospital Medicine. 2006;67(2):60-1.
  10. Schuch FB, et al. Exercise as a treatment for depression: A meta-analysis adjusting for publication bias. Journal of Psychiatric Research. 2016;77:42-51.
  11. Morres ID, et al. Aerobic exercise for adult patients with major depressive disorder in mental health services: A systematic review and meta-analysis. Depression and Anxiety. 2019;36(1):39-53.
  12. Krogh J, et al. Exercise for patients with major depression: a systematic review with meta-analysis and trial sequential analysis. BMJ Open. 2017;7(9):e014820.
  13. North TC, McCullagh P, Tran ZV. Effect of Exercise on Depression. Exercise and Sport Sciences Reviews. 1990;18(1).
  14. Honey EP. I shrunk the pooled SMD! Guide to critical appraisal of systematic reviews and meta-analyses using the Cochrane review on exercise for depression as example. Ment Health Phys Act. 2015;8:21-36.
  15. Yu Q, et al. Comparative Effectiveness of Multiple Exercise Interventions in the Treatment of Mental Health Disorders: A Systematic Review and Network Meta-Analysis. Sports Medicine - Open. 2022;8(1):135.
  16. Singh B, et al. Effectiveness of physical activity interventions for improving depression, anxiety, and distress: an overview of systematic reviews. British Journal of Sports Medicine. 2023;57(18):1203-9.
  17. Stroup DF, et al. Meta-analysis of observational studies in epidemiology: a proposal for reporting. JAMA. 2000;283(15):2008-12.
  18. Page MJ, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.
  19. Liberati A, et al. The PRISMA statement for reporting systematic reviews and meta-analyses of studies that evaluate health care interventions: explanation and elaboration. J Clin Epidemiol. 2009;62(10):e1-e34.
  20. Gupta S, et al. Efficacy of generic oral directly acting agents in patients with hepatitis C virus infection. J Viral Hepat. 2018;25(7):771-8.
  21. Higgins J, et al. eds. Cochrane Handbook for Systematic Reviews of Interventions. 2nd ed. Chichester, UK: Wiley & Sons; 2019.
  22. DerSimonian R, Laird N. Meta-analysis in clinical trials. Controlled Clinical Trials. 1986;7(3):177-88.
  23. Higgins JP, Thompson SG, Deeks JJ, Altman DG. Measuring inconsistency in meta-analyses. Bmj. 2003;327(7414):557-60.
  24. Nimmo MA, McLean D, Mutrie N, McKenzie S. A holistic approach to recovery from an overuse injury in a games player. British Journal of Sports Medicine. 1986;20(3):103-6.
  25. Doyne EJ, et al. Running versus weight lifting in the treatment of depression. Journal of consulting and clinical psychology. 1987;55(5):748.
  26. McNeil JK, LeBlanc EM, Joyner M. The effect of exercise on depressive symptoms in the moderately depressed elderly. Psychology and aging. 1991;6(3):487.
  27. Veale D, et al. Aerobic Exercise in the Adjunctive Treatment of Depression: A Randomized Controlled Trial. Journal of the Royal Society of Medicine. 1992;85(9):541-4.
  28. Singh NA, Clements KM, Fiatarone MA. A Randomized Controlled Trial of Progressive Resistance Training in Depressed Elders. The Journals of Gerontology: Series A. 1997;52A(1):M27-M35.
  29. Mather AS, et al. Effects of exercise on depressive symptoms in older adults with poorly responsive depressive disorder: Randomized controlled trial. British Journal of Psychiatry. 2002;180(5):411-5.
  30. Nabkasorn C, et al. Effects of physical exercise on depression, neuroendocrine stress hormones, and physiological fitness in adolescent females with depressive symptoms. European Journal of Public Health. 2005;16(2):179-84.
  31. Singh NA, et al. A Randomized Controlled Trial of High Versus Low Intensity Weight Training Versus General Practitioner Care for Clinical Depression in Older Adults. The Journals of Gerontology: Series A. 2005;60(6):768-76.
  32. Sims J, et al. Exploring the feasibility of a community-based strength training program for older people with depressive symptoms and its impact on depressive symptoms. BMC Geriatrics. 2006;6(1):18.
  33. Brenes GA, et al. Treatment of minor depression in older adults: a pilot study comparing sertraline and exercise. Aging Ment Health. 2007;11(1):61-8.
  34. Blumenthal JA, et al. Exercise and pharmacotherapy in the treatment of major depressive disorder. Psychosomatic medicine. 2007;69(7):587-96.
  35. Pilu A, et al. Efficacy of physical activity in the adjunctive treatment of major depressive disorders: preliminary results. Clinical Practice and Epidemiology in Mental Health. 2007;3(1):8.
  36. Vieira JLL, Porcu M, Rocha PGMd. A prática de exercícios físicos regulares como terapia complementar ao tratamento de mulheres com depressão. Jornal Brasileiro de Psiquiatria. 2007;56:23-8.
  37. Williams CL, Tappen RM. Exercise training for depressed older adults with Alzheimer's disease. Aging & Mental Health. 2008;12(1):72-80.
  38. Sims J, et al. Regenerate: assessing the feasibility of a strength-training program to enhance the physical and mental health of chronic post stroke patients with depression. International Journal of Geriatric Psychiatry. 2009;24(1):76-83.
  39. Shahidi M, et al. Laughter yoga versus group exercise program in elderly depressed women: a randomized controlled trial. International Journal of Geriatric Psychiatry. 2011;26(3):322-7.
  40. Mota-Pereira J, et al. Moderate exercise improves depression parameters in treatment-resistant patients with major depressive disorder. Journal of Psychiatric Research. 2011;45(8):1005-11.
  41. Prakhinkit S, Suppapitiporn S, Tanaka H, Suksom D. Effects of Buddhism Walking Meditation on Depression, Functional Fitness, and Endothelium-Dependent Vasodilation in Depressed Elderly. The Journal of Alternative and Complementary Medicine. 2013;20(5):411-6.
  42. Legrand FD. Effects of Exercise on Physical Self-Concept, Global Self-Esteem, and Depression in Women of Low Socioeconomic Status With Elevated Depressive Symptoms. Journal of Sport and Exercise Psychology. 2014;36(4):357-65.
  43. Pfaff JJ, et al. ACTIVEDEP: a randomised, controlled trial of a home-based exercise intervention to alleviate depression in middle-aged and older adults. British Journal of Sports Medicine. 2014;48(3):226-32.
  44. Danielsson L, et al. Exercise or basic body awareness therapy as add-on treatment for major depression: A controlled study. Journal of Affective Disorders. 2014;168:98-106.
  45. Oertel-Knöchel V, et al. Effects of aerobic exercise on cognitive performance and individual psychopathology in depressive and schizophrenia patients. European Archives of Psychiatry and Clinical Neuroscience. 2014;264(7):589-604.
  46. Hallgren M, et al. Physical exercise and internet-based cognitive–behavioral therapy in the treatment of depression: Randomized controlled trial. British Journal of Psychiatry. 2015;207(3):227-34.
  47. Carneiro LSF, et al. Effects of structured exercise and pharmacotherapy vs. pharmacotherapy for adults with depressive symptoms: A randomized clinical trial. Journal of Psychiatric Research. 2015;71:48-55.
  48. Doose M, et al. Self-selected intensity exercise in the treatment of major depression: A pragmatic RCT. International Journal of Psychiatry in Clinical Practice. 2015;19(4):266-75.
  49. Schuch FB, et al. Exercise and severe major depression: Effect on symptom severity and quality of life at discharge in an inpatient cohort. Journal of Psychiatric Research. 2015;61:25-32.
  50. Gao L, Zhang L, Qi H, Petridis L. Middle-aged female depression in perimenopausal period and square dance intervention. Psychiatria Danubina. 2016;28(4):372-8.
  51. Schneider KL, et al. Feasibility of Pairing Behavioral Activation With Exercise for Women With Type 2 Diabetes and Depression: The Get It Study Pilot Randomized Controlled Trial. Behavior Therapy. 2016;47(2):198-212.
  52. Lok N, Lok S, Canbaz M. The effect of physical activity on depressive symptoms and quality of life among elderly nursing home residents: Randomized controlled trial. Archives of Gerontology and Geriatrics. 2017;70:92-8.
  53. Cheung LK, Lee S. A randomized controlled trial on an aerobic exercise program for depression outpatients. Sport Sciences for Health. 2018;14(1):173-81.
  54. Roy A, Govindan R, Muralidharan K. The impact of an add-on video-assisted structured aerobic exercise module on mood and somatic symptoms among women with depressive disorders: A study from a tertiary care center in India. Asian Journal of Psychiatry. 2018;32:118-22.
  55. Abdelbasset WK, et al. Examining the impacts of 12 weeks of low to moderate-intensity aerobic exercise on depression status in patients with systolic congestive heart failure- A randomized controlled study. Clinics. 2019;74:e1017.
  56. Makizako H, et al. Exercise and Horticultural Programs for Older Adults with Depressive Symptoms and Memory Problems: A Randomized Controlled Trial. Journal of Clinical Medicine. 2020;9(1):99.
  57. La Rocque CL, et al. Randomized controlled trial of Bikram yoga and aerobic exercise for depression in women: Efficacy and stress-based mechanisms. Journal of Affective Disorders. 2021;280:457-66.
  58. Chau RMW, et al. Effectiveness of a structured physical rehabilitation program on the physical fitness, mental health, and pain for Chinese patients with major depressive disorders in Hong Kong – a randomized controlled trial with 9-month follow-up outcomes. Disability and Rehabilitation. 2022;44(8):1294-304.
  59. Nicholas M, Nsibambi An, Ojuka E, Maghanga M. Twelve Weeks Aerobic Exercise Improves Anxiety and Depression in HIV Positive Clients on Art in Uganda. International Journal of Sport Exercise and Training Sciences - IJSETS. 2024;10(4):288-98.
  60. Woolf C, et al. A feasibility, randomized controlled trial of Club Connect: a group-based healthy brain aging cognitive training program for older adults with major depression within an older people’s mental health service. BMC Psychiatry. 2024;24(1):208.
  61. James Vibin A, et al. Effect of Integrated Yoga as an add-on therapy in adults with clinical depression – A randomized controlled trial. International Journal of Social Psychiatry. 2024;70(4):709-19.
  62. Chang Q, et al. Effects of varying exercise intensities on muscle strength and depressive symptoms in Chinese adolescents: A 12-week randomized controlled trial. PLOS ONE. 2025;20(11):e0336894.
  63. Wei X, et al. Efficacy of a Combined Aerobic Exercise and Mindfulness Intervention on Depressive Symptoms: A Pilot Randomized Controlled Trial. Research on Social Work Practice. 2025;10497315251316844.
  64. Ross R, Janssen I, Tremblay MS. Public health importance of light intensity physical activity. Journal of Sport and Health Science. 2024;13(5):674-5.
  65. Maltagliati S, et al. Effort minimization: A permanent, dynamic, and surmountable influence on physical activity. J Sport Health Sci. 2025;14:100971.
  66. Cuijpers P, et al. Psychotherapy for Depression Across Different Age Groups: A Systematic Review and Meta-analysis. JAMA Psychiatry. 2020;77(7):694-702.
  67. Cipriani A, et al. Comparative efficacy and acceptability of 21 antidepressant drugs for the acute treatment of adults with major depressive disorder: a systematic review and network meta-analysis. The Lancet. 2018;391(10128):1357-66.
  68. Stubbs B, et al. Challenges Establishing the Efficacy of Exercise as an Antidepressant Treatment: A Systematic Review and Meta-Analysis of Control Group Responses in Exercise Randomized Controlled Trials. Sports Medicine. 2016;46(5):699-713.
  69. Gupta SK. Intention-to-treat concept: A review. Perspectives in Clinical Research. 2011;2(3).
  70. Cullen W, Gulati G, Kelly BD. Mental health in the COVID-19 pandemic. QJM: An International Journal of Medicine. 2020;113(5):311-2.
  71. Hardcastle SJ, et al. A randomized controlled trial of Promoting Physical Activity in Regional and Remote Cancer Survivors (PPARCS). J Sport Health Sci. 2024;13(1):81-9.
  72. Herold F, et al. Alexa, let's train now! - A systematic review and classification approach to digital and home-based physical training interventions aiming to support healthy cognitive aging. J Sport Health Sci. 2024;13(1):30-46.
  73. Liu Y, et al. The efficacy of exercise interventions on depressive symptoms and cognitive function in adults with depression: An umbrella review. Journal of Affective Disorders. 2025;368:779-88.
  74. Munro NR, et al. Effect of exercise on depression and anxiety symptoms: systematic umbrella review with meta-meta-analysis. British Journal of Sports Medicine. 2026;60(8):590-9.

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

Randomized Controlled TrialsRemission RatesStandardized Mean DifferenceOdds RatioSupervised ExerciseControl GroupsGrading Of Recommendations

Related Articles