Study design and registration
This systematic review was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 statement and the Cochrane Handbook for Systematic Reviews of Interventions. The completed PRISMA 2020 checklist is provided in Supplementary File 1. The review protocol was not registered prospectively10,11.
Search strategy and timeframe
A comprehensive literature search was performed in the following electronic databases: Medline (via PubMed), Embase (via Elsevier), and the Cochrane Central Register of Controlled Trials (CENTRAL). The search period extended from database inception (January 1, 1946, for Medline; January 1, 1947, for Embase; and January 1, 1995, for CENTRAL) through March 31, 2026. The search strategies combined controlled vocabulary (MeSH terms for Medline and CENTRAL; Emtree terms for Embase) and free-text keywords related to two core concepts: (1) spinal cord injury (including "spinal cord injury," "SCI," "spinal trauma," "spine fracture," "cervical fracture," and "cervical trauma") and (2) magnetic resonance imaging (including "MRI," "magnetic resonance imaging," "diffusion tensor imaging," "DTI," "susceptibility weighted imaging," and "SWI"). No language restrictions were applied initially, but the final inclusion was limited to English-language articles12.
Eligibility criteria
Studies were included if they met the following criteria: (1) human subjects; (2) adults aged 16 years or older with acute spinal cord injury (within 7 days of injury); (3) MRI performed within 7 days of injury; (4) study addressed one or more of the six prespecified key questions (diagnostic accuracy, frequency of findings, influence on decision-making, optimal timing, safety, or outcomes); (5) English language; and (6) original research, including randomized controlled trials, prospective or retrospective cohort studies, case-control studies, and case series with 10 or more patients. Exclusion criteria were: (1) pediatric populations (age <16 years); (2) MRI performed solely for prognostic purposes without a clinical decision-making context; (3) review articles, opinion pieces, editorials, case reports, or case series with fewer than 10 patients; and (4) animal or biomechanical studies.
Screening process
Two authors (W.Z. and J.S.) independently screened all titles and abstracts retrieved from the database searches using a standardized screening form. Any citation deemed potentially relevant by either reviewer advanced to full-text review. The same two authors independently assessed the full text of each potentially eligible article against the prespecified inclusion and exclusion criteria. Disagreements at either the title/abstract or full-text screening stage were resolved through discussion and consensus; if consensus could not be reached, a third author (J.H.) served as arbitrator and made the final determination. The screening process was managed using Covidence systematic review software. Inter-rater agreement at the full-text screening stage was calculated using Cohen's kappa coefficient, which was 0.89 (95% CI, 0.84–0.94), indicating near-perfect agreement.
Data extraction and validation
Data extraction was performed independently by two authors (W.Z. and K.W.) using a standardized, piloted data extraction template developed in Microsoft Excel. The template included the following fields: study characteristics (first author, year of publication, country, study design, sample size), population demographics (age, sex, injury level, ASIA Impairment Scale grade at presentation), MRI protocol (field strength, sequences acquired, timing post-injury), findings relevant to each key question (diagnostic accuracy metrics, frequencies of abnormal findings, decision-altering events, timing data, adverse events, and outcome measures), and reported effect sizes (odds ratios, mean differences, or proportions with confidence intervals). After independent extraction, the two authors compared their extracted data. Discrepancies were resolved by re-reviewing the original article and discussing until consensus was achieved; if disagreement persisted, a third author (H.Y.) adjudicated. No automated data extraction tools were used. For studies with missing or unclear data, corresponding authors were contacted via email up to two times over a four-week period; if no response was received, the available data were reported as presented.
Risk of bias and quality assessment
Two authors (J.S. and K.W.) independently assessed the risk of bias for each included study using the National Heart, Lung, and Blood Institute (NHLBI) Quality Assessment Tool for Observational Cohort and Cross-Sectional Studies. Each study was rated as "good" (low risk of bias, with valid results that are unlikely to change with further research), "fair" (moderate risk of bias, with some limitations but not sufficient to invalidate the results), or "poor" (high risk of bias, with significant methodological flaws). Disagreements in quality ratings were resolved through consensus adjudicated by a third author (J.H.). Studies rated as "poor" were not excluded a priori but were subjected to sensitivity analyses to assess their impact on pooled effect estimates.
Statistical analysis and data synthesis
All statistical analyses were performed using R version 4.2.2. Due to anticipated clinical and methodological heterogeneity across studies, all meta-analyses were performed using random-effects models with the DerSimonian-Laird estimator for the between-study variance (τ2). For frequency data (proportions), pooled estimates with 95% confidence intervals were calculated using the inverse-variance method, with the Freeman-Tukey double arcsine transformation to stabilize variances. For comparative outcome data, pooled odds ratios with 95% confidence intervals were calculated using the Mantel-Haenszel method. For continuous outcomes (e.g., length of stay, motor scores), pooled mean differences with 95% confidence intervals were calculated using the inverse variance method.
Between-study heterogeneity was assessed using the I2 statistic and Cochran's Q test, with I2 values of 25%, 50%, and 75% interpreted as low, moderate, and high heterogeneity, respectively. Given the anticipated high heterogeneity (I2 potentially > 90%) due to variations in injury severity, MRI protocols, timing of imaging, and study designs, the following a priori subgroup and sensitivity analyses were planned and executed to explore potential sources of heterogeneity.
Subgroup analyses
For outcomes with I2 > 75%, subgroup analyses were performed based on the following prespecified variables: (1) injury level (cervical vs. thoracolumbar); (2) presence of fracture on CT (fracture/dislocation vs. SCIWORA); (3) MRI field strength (1.5T vs. 3T); (4) MRI sequences used (conventional vs. advanced sequences including STIR, DTI, SWI); (5) study design (prospective vs. retrospective); and (6) risk of bias rating (good vs. fair vs. poor). Subgroup differences were assessed using mixed-effects meta-regression, with p < 0.10 considered statistically significant for interaction due to the exploratory nature of these analyses.
Sensitivity analyses
To assess the robustness of pooled estimates in the presence of high heterogeneity, sensitivity analyses were performed, including restriction to studies rated as "good" quality (low risk of bias). Publication bias was evaluated using funnel plots for outcomes with 10 or more studies and statistically using Egger's linear regression test.
Reporting of heterogeneity
For all pooled estimates, the I2 statistic and its 95% confidence interval (where calculable) are reported. When I2 exceeded 75%, the pooled estimate is presented with a cautionary note, and the results of subgroup and sensitivity analyses are reported in the text to guide interpretation. When subgroup analyses fail to explain substantial heterogeneity (residual I2 > 75% after subgrouping), the pooled estimate is reported as an average of highly variable effects, and readers are advised to interpret it with appropriate caution.
Publication bias was evaluated visually using funnel plots for outcomes with 10 or more studies and statistically using Egger's linear regression test. All p-values were two-sided, with statistical significance set at p < 0.05. For all analyses, 95% confidence intervals are reported in accordance with standard scientific formatting (e.g., 95% CI, 1.32–2.41).