$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
The representative results show the outputs generated by the 16 week competition-integrated Strategic Management course improvement protocol. Results are organized by implementation fidelity, instructional quality, student-level outcomes, team-level outcomes, and predefined warning patterns. Because the data came from routine educational cohorts with a historical comparison group rather than randomized assignment, all results should be interpreted as context-specific implementation and evaluation outputs, not causal evidence of effectiveness.
Implementation fidelity and instructional quality
The available educational-record dataset included 88 students in the historical comparison cohort and 92 students in the protocol cohort. The team-level dataset included 22 comparison-cohort teams and 23 protocol-cohort teams. After applying the paired pre-post knowledge-test rule, 79 comparison-cohort students and 84 protocol-cohort students were retained for knowledge-gain analysis.
Implementation fidelity was checked before learning outcomes were interpreted. The protocol cohort completed 15 of 16 planned weekly activities (93.8%), whereas the comparison cohort completed 13 of 16 (81.3%). The enterprise or structured case session was completed in the protocol cohort but was not part of the full historical comparison structure. In the protocol cohort, 88 of 92 students completed all three competition tasks, corresponding to 95.7%. In the comparison cohort, 18 of 88 students had complete records for three comparable project-based assessments, corresponding to 20.5%; these records were retained for descriptive comparison only because the historical course did not implement the full competition sequence. Feedback was returned within 7 calendar days for 243 of 268 protocol-cohort submissions, corresponding to 90.7%, compared with 49 of 66 comparison-cohort submissions, corresponding to 74.2%. Mean feedback turnaround was 5.8 ± 1.9 days and 8.4 ± 2.7 days, respectively. Observation completion reached 81.3% and 62.5%. The protocol-cohort two-rater ICC for overlapping subjective scoring was 0.82, exceeding the predefined 0.75 threshold. Overall, the protocol cohort achieved 5 of 5 core implementation-fidelity indicators, whereas the comparison cohort achieved 2 of 5.
Instructional quality indicators were higher in the protocol cohort. Classroom interaction was 3.70 ± 0.44 in the protocol cohort and 3.15 ± 0.48 in the comparison cohort. Case-design completeness was 4.08 ± 0.36 and 3.23 ± 0.42. Task-rubric alignment was 4.01 ± 0.39 and 2.30 ± 0.51. Student engagement increased by 0.42 ± 0.41 points in the protocol cohort, exceeding the predefined 0.30-point signal threshold, whereas the comparison cohort increased by 0.07 ± 0.34 points. Attendance was similar between cohorts: 89.6% ± 5.8% and 89.1% ± 6.2%. Table 5 summarizes these indicators.
| Indicator | Data source | Predefined threshold | Historical comparison cohort | Protocol cohort | Interpretation |
| Planned weekly activity completion | Course implementation log | ≥85% | 13/16 weeks, 81.3% | 15/16 weeks, 93.8% | Achieved in protocol cohort |
| Enterprise or structured case-session implementation | Enterprise engagement log | ≥1 implemented session | 0/1 session, 0.0% | 1/1 session, 100.0% | Achieved in protocol cohort |
| Complete task or mapped project-record completion | Submission records | ≥85% | 18/88 students, 20.5% | 88/92 students, 95.7% | Achieved in protocol cohort; historical records descriptive only |
| Feedback returned within 7 days | Feedback log | ≥80% | 49/66 submissions, 74.2% | 243/268 submissions, 90.7% | Achieved in protocol cohort |
| Observation checklist completion | Observation forms | ≥75% | 10/16 sessions, 62.5% | 13/16 sessions, 81.3% | Achieved in protocol cohort |
| Two-rater agreement | Rater scores | ICC ≥0.75 | Not required | ICC = 0.82 | Achieved in protocol cohort |
| Feedback turnaround time | Feedback log | Mean ≤7 days preferred | 8.4 ± 2.7 days | 5.8 ± 1.9 days | Within preferred range in protocol cohort |
| Classroom interaction | Observation checklist | Mean ≥3.50 or increase ≥0.30 | 3.15 ± 0.48 | 3.70 ± 0.44 | Higher in protocol cohort |
| Case-design completeness | Case-design checklist | Mean ≥4.00 preferred | 3.23 ± 0.42 | 4.08 ± 0.36 | Higher in protocol cohort |
| Task-rubric alignment | Alignment checklist | Mean ≥4.00 | 2.30 ± 0.51 | 4.01 ± 0.39 | Achieved in protocol cohort |
| Student engagement gain | Attendance, logs, survey items | Increase ≥0.30 points | +0.07 ± 0.34 | +0.42 ± 0.41 | Achieved in protocol cohort |
| Attendance rate | Attendance record | Descriptive | 89.1% ± 6.2% | 89.6% ± 5.8% | Similar across cohorts |
| Core implementation-fidelity summary | Implementation indicators | Most core thresholds achieved | 2/5 thresholds achieved | 5/5 thresholds achieved | Protocol implemented as intended |
Table 5: Representative implementation fidelity and instructional quality indicators. This table reports representative implementation fidelity and instructional quality outputs for the historical comparison cohort and protocol cohort and indicates whether the predefined operational thresholds were achieved.
Figure 2 presents the implementation fidelity and instructional quality profile. Figure 2A shows planned weekly activity completion. Figure 2B shows three-task sequence completion. Figure 2C shows feedback returned within 7 calendar days. Figure 2D shows classroom interaction, case design completeness, task rubric alignment, and student engagement gains. Figure 2E shows observation completion and enterprise or structured case-session completion. Percentage-based fidelity measures and scale-based instructional indicators should be interpreted within their own measurement scales.

Figure 2: Implementation fidelity and instructional quality profile. This figure summarizes implementation fidelity and instructional quality indicators for the historical comparison cohort and the protocol cohort. (A) Planned weekly activity completion. (B) Complete task or mapped project-record completion, distinguishing the protocol cohort’s full three-task sequence from mapped project records in the historical comparison cohort. (C) Feedback returned within 7 calendar days, with feedback-turnaround time shown as supporting information. (D) Instructional quality indicators, including classroom interaction, case-design completeness, task-rubric alignment, and student engagement gain. (E) Observation checklist completion and enterprise or structured case-session implementation. Please click here to view a larger version of this figure.
Student-level learning outcomes
Student-level outcomes were examined separately from team-level competition outcomes. Baseline knowledge scores were 64.18 ± 8.74 in the protocol cohort and 62.37 ± 8.91 in the comparison cohort, with a small baseline difference of 1.81 points (d = 0.21). Post-course knowledge scores were 71.08 ± 8.58 and 66.41 ± 9.06, with a difference of 4.67 points (d = 0.53). Knowledge gain was 6.90 ± 5.41 points and 4.04 ± 5.26 points, with a difference of 2.86 points (d = 0.54). The protocol cohort exceeded the predefined 5-point knowledge-gain signal.
Late-semester individual case-analysis performance was 82.31 ± 7.06 in the protocol cohort and 76.42 ± 7.38 in the comparison cohort, with a difference of 5.89 points (d = 0.82). This outcome was interpreted as a postcourse performance indicator rather than a paired gain because no true baseline case-analysis assessment was administered. Satisfaction was 3.66 ± 0.58 and 3.34 ± 0.62 (d = 0.53). Perceived practical relevance was 3.91 ± 0.55 and 3.41 ± 0.59 (d = 0.88). Competition pressure was higher in the protocol cohort, 3.12 ± 0.76, compared with 2.48 ± 0.71 (d = 0.87), but high workload pressure was reported by 13 of 92 protocol-cohort students, corresponding to 14.1%, below the predefined 20% warning threshold. Final course score was treated as a secondary educational outcome and was 87.20 ± 6.10 in the protocol cohort and 84.80 ± 6.70 in the comparison cohort (d = 0.37). Table 6 summarizes student-level and team-level outcomes.
| Outcome | Level | Historical comparison cohort | Protocol cohort | Difference | Effect size | Interpretation |
| Available educational-record sample | Student | n = 88 | n = 92 | — | — | Usable de-identified course records |
| Paired knowledge-analysis sample | Student | n = 79 | n = 84 | — | — | Students with baseline and post-course knowledge scores |
| Baseline knowledge score | Student | 62.37 ± 8.91 | 64.18 ± 8.74 | 1.81 | 0.21 | Small baseline difference |
| Post-course knowledge score | Student | 66.41 ± 9.06 | 71.08 ± 8.58 | 4.67 | 0.53 | Higher in protocol cohort |
| Knowledge gain | Student | 4.04 ± 5.26 | 6.90 ± 5.41 | 2.86 | 0.54 | Protocol cohort exceeded the ≥5-point signal |
| Individual case-analysis score | Student | 76.42 ± 7.38 | 82.31 ± 7.06 | 5.89 | 0.82 | Exceeded the ≥3-point case-analysis signal |
| Student engagement gain | Student | +0.07 ± 0.34 | +0.42 ± 0.41 | +0.35 | 0.93 | Exceeded the ≥0.30-point signal |
| Satisfaction score | Student | 3.34 ± 0.62 | 3.66 ± 0.58 | 0.32 | 0.53 | Higher in protocol cohort |
| Perceived practical relevance | Student | 3.41 ± 0.59 | 3.91 ± 0.55 | 0.5 | 0.88 | Higher in protocol cohort |
| Competition pressure score | Student | 2.48 ± 0.71 | 3.12 ± 0.76 | 0.64 | 0.87 | Higher pressure in protocol cohort |
| High workload pressure | Student | Not assessed as competition pressure | 13/92, 14.1% | — | — | Below the 20% warning threshold |
| Final course score | Student | 84.80 ± 6.70 | 87.20 ± 6.10 | 2.4 | 0.37 | Secondary educational outcome |
| Number of teams | Team | n = 22 | n = 23 | — | — | Team-level analyses reported separately |
| Decision-round exposure | Team | 0.77 ± 0.81 rounds | 5.09 ± 0.79 rounds | 4.32 | Not estimated | Implementation exposure indicator |
| Task 1 or mapped diagnosis score | Team | 69.84 ± 6.45 | 72.18 ± 6.28 | 2.34 | 0.37 | Modest team-level difference |
| Task 2 or mapped decision score | Team | 71.09 ± 6.12 | 73.52 ± 5.91 | 2.43 | 0.4 | Modest team-level difference |
| Task 3 or mapped final-defense score | Team | 72.71 ± 6.36 | 77.20 ± 6.04 | 4.49 | 0.72 | Largest team-level difference |
| Final team score | Team | 71.21 ± 5.81 | 74.30 ± 5.52 | 3.09 | 0.55 | Moderate team-level difference |
| Strategic diagnosis dimension | Team | 14.60 ± 1.80 | 15.60 ± 1.70 | 1 | — | Final-defense rubric dimension |
| Strategic choice dimension | Team | 17.90 ± 2.20 | 19.10 ± 2.10 | 1.2 | — | Final-defense rubric dimension |
| Implementation feasibility dimension | Team | 14.40 ± 1.90 | 15.50 ± 1.80 | 1.1 | — | Final-defense rubric dimension |
| Innovation dimension | Team | 10.80 ± 1.60 | 11.50 ± 1.50 | 0.7 | — | Final-defense rubric dimension |
| Evidence use dimension | Team | 7.10 ± 1.00 | 7.50 ± 0.90 | 0.4 | — | Final-defense rubric dimension |
| Oral defense dimension | Team | 7.91 ± 1.10 | 8.00 ± 1.00 | 0.09 | — | Final-defense rubric dimension |
| Peer contribution mean | Team | 3.72 ± 0.41 | 3.91 ± 0.36 | 0.19 | 0.49 | No major team-process failure |
| Teams with contribution imbalance | Team | 4/22, 18.2% | 3/23, 13.0% | −5.2 percentage points | — | No systematic imbalance pattern |
| Award distribution | Team | Not applicable | 10 none; 8 third; 2 second; 3 first | — | — | Protocol cohort only |
| Score-inflation warning pattern | Team | Not applicable | 2/23 teams, 8.7% | — | — | Flagged for troubleshooting |
Table 6:Representative student-level and team-level outcomes. This table reports representative student-level and team-level outcomes, including available educational-record sample, paired knowledge-analysis sample, knowledge scores, knowledge gain, individual case-analysis score, engagement, satisfaction, perceived practical relevance, workload pressure, final course score, team task scores, final-defense rubric-dimension scores, peer contribution, award distribution, and warning indicators.
Figure 3 presents student-level outcomes with uncertainty estimates where applicable. Figure 3A shows baseline and post-course knowledge scores. Figure 3B shows knowledge gain. Figure 3C shows individual case-analysis performance. Figure 3D shows engagement gain, satisfaction, and perceived practical relevance. Figure 3E shows the final course score as a secondary outcome.

Figure 3: Student-level representative learning outcomes. This figure presents individual-level outcomes separately from team-level competition outcomes. (A) Baseline and post-course Strategic Management knowledge scores. (B) Knowledge gain from baseline to post-course assessment. (C) Late-semester individual case-analysis score. (D) Engagement gain and survey perception indicators, including satisfaction and perceived practical relevance, are shown with separate scales where required. (E) Final course score as a secondary educational outcome and protocol-cohort high workload pressure relative to the predefined warning threshold. Values are presented as mean ± SD where ± values are shown, and error bars represent SD. In (E), the dashed reference line indicates the predefined high-workload-pressure warning threshold of 20%. Please click here to view a larger version of this figure.
Team-level competition outcomes
Team-level outcomes were analyzed separately from student-level outcomes. For descriptive reporting, comparison-cohort case-project records were mapped to comparable assessment points when available, but these records did not represent completion of the full competition-integrated sequence. Decision-round exposure was reported as an implementation exposure indicator rather than a learning-effect estimate. The protocol cohort completed 5.09 ± 0.79 decision rounds, whereas the comparison cohort completed 0.77 ± 0.81 mapped rounds.
In the protocol cohort, the mean Task 1 score was 72.18 ± 6.28, the mean Task 2 score was 73.52 ± 5.91, and the mean Task 3 score was 77.20 ± 6.04. The corresponding mapped comparison-cohort scores were 69.84 ± 6.45, 71.09 ± 6.12, and 72.71 ± 6.36. The final team score was 74.30 ± 5.52 in the protocol cohort and 71.21 ± 5.81 in the comparison cohort. Final-defense or mapped final-project rubric dimensions were as follows: strategic diagnosis, 15.60 ± 1.70 and 14.60 ± 1.80; strategic choice, 19.10 ± 2.10 and 17.90 ± 2.20; implementation feasibility, 15.50 ± 1.80 and 14.40 ± 1.90; innovation, 11.50 ± 1.50 and 10.80 ± 1.60; evidence use, 7.50 ± 0.90 and 7.10 ± 1.00; and oral defense, 8.00 ± 1.00 and 7.91 ± 1.10, respectively.
Award distribution was reported for the protocol cohort only because the comparison cohort did not use the full award-based competition structure. In the protocol cohort, 10 teams received no award, 8 received third prize, 2 received second prize, and 3 received first prize. Mean peer contribution was 3.91 ± 0.36 in the protocol cohort and 3.72 ± 0.41 in the comparison cohort. Contribution imbalance was identified in 3 of 23 protocol-cohort teams and 4 of 22 comparison-cohort teams. A score-inflation warning pattern was identified in 2 of 23 protocol-cohort teams, corresponding to 8.7%.
Figure 4 presents team-level outcomes and warning patterns. Figure 4A shows Task 1, Task 2, and Task 3 team scores. Figure 4B shows final-defense rubric-dimension scores. Figure 4C shows protocol-cohort award distribution only. Figure 4D shows score-inflation warning cases. Figure 4E shows peer contribution balance and contribution-imbalance counts.

Figure 4: Team-level competition outcomes and warning patterns. This figure presents team-level outcomes separately from student-level learning outcomes. (A) Task-stage team scores for Task 1 strategic diagnosis, Task 2 strategic decision, and Task 3 final defense, with mapped project scores shown for the historical comparison cohort where available. (B) Final-defense rubric-dimension scores, including strategic diagnosis, strategic choice, implementation feasibility, innovation, evidence use, and oral defense. (C) Protocol-cohort award distribution only, because the historical comparison cohort did not use the full award-based competition structure. (D) Score-inflation warning pattern, defined as high final-defense performance without supporting diagnostic strength or individual learning evidence. (E) Peer contribution score and contribution-imbalance counts are displayed as separate team-level indicators. Please click here to view a larger version of this figure.
Suboptimal patterns and troubleshooting indicators
Two of 23 protocol-cohort teams showed a predefined score-inflation warning pattern. These teams had relatively high final defense scores but weaker Task 1 diagnostic scores and below-average individual knowledge gains. This pattern was flagged for follow-up in the next course iteration rather than interpreted as evidence of learning improvement. No major ceiling-effect warning was observed because baseline knowledge scores were below 85 out of 100, and baseline satisfaction was below 4.50 out of 5. Attendance was similar between cohorts, and no systematic team contribution imbalance was observed, although three protocol-cohort teams required monitoring.
Overall, the representative results showed that the protocol generated interpretable outputs at three levels: implementation fidelity, student-level learning evidence, and team-level competition performance. The warning cases supported the use of the protocol for course monitoring, troubleshooting, and subsequent course refinement.
The de-identified dataset supporting this study has been deposited in figshare.com and is available at https://doi.org/10.6084/m9.figshare.33094931. The dataset includes student-level records, team-level competition records, implementation indicators, participant-flow records, missing-data audit records, and rater-score records. Student identifiers, consent records, linkage files, and identifiable free-text comments are not included.
Supplementary Table S1: Rater training, scoring agreement, and score reconciliation procedure. This table provides the rater assignment, independence check, rubric orientation, calibration scoring, scoring agreement check, score-gap rule, reconciliation procedure, data-entry audit, score lock, and archive procedure.Please click here to download this file.
Supplementary Table S2: Troubleshooting matrix for protocol deviations and suboptimal implementation. This table lists protocol deviations and suboptimal implementation patterns, including trigger rules, corrective actions, required records, and interpretation notes.Please click here to download this file.
Supplementary Table S3: Dataset structure and variable dictionary for the de-identified educational dataset. This table defines student-level, team-level, and class-session-level variables, including variable codes, descriptions, data types, coding ranges, data sources, and missing-data rules.Please click here to download this file.
Supplementary File 1: Operational templates, assessment materials, and de-identified dataset workbook. This supplementary file provides the operational templates, assessment materials, scoring forms, survey items, case-design check, missing-data rules, de-identified dataset workbook structure, cleaning log, analysis summary template, and version-control record used to support reproducibility.Please click here to download this file.