Leakage-conscious machine learning for supplied injury-risk label classification and longitudinal injury forecasting: a multi-dataset evaluation
BMC Sports Science, Medicine and Rehabilitation, cilt.2026, sa.1, ss.1-30, 2026 (Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 2026 Sayı: 1
- Basım Tarihi: 2026
- Doi Numarası: 10.1186/s13102-026-02109-6
- Dergi Adı: BMC Sports Science, Medicine and Rehabilitation
- Derginin Tarandığı İndeksler: Scopus
- Sayfa Sayıları: ss.1-30
- Kocaeli Üniversitesi Adresli: Evet
Özet
Background
Machine learning is increasingly used to model sports-related injury risk, but favorable performance may be misleading when outcome labels are externally assigned, proxy-driven, highly separable, or affected by outcome-proximal information. This study evaluated model credibility across heterogeneous injury-related settings, emphasizing leakage control, baseline comparison, calibration, uncertainty, temporal validation, and explainability.
Methods
The primary analysis used SoccerMon, a longitudinal athlete-monitoring dataset, to forecast the start of a reported location-specific injury episode within 14 days using internal season-to-season temporal validation in the same athlete panel. Two externally supplied, pre-labelled datasets—the Personalized Sports Health Dataset and Biomechanical Analysis for Injury Prevention—were retained as supporting methodological stress tests for three-class classification of supplied low-, medium-, and high-risk labels. Their participant provenance, real-versus-synthetic status, and target-generation procedures could not be independently verified. Evaluation included structure-appropriate splitting, repeated cross-validation, athlete-cluster uncertainty, calibration, missing-data and proxy-removal sensitivity analyses, event-level assessment, and explainability diagnostics.
Results
In the 2021 SoccerMon internal temporal test period (50 athletes; 12 injury episodes; 17,535 player-days), the class-weighted focal Logistic Regression achieved ROC-AUC 0.7558 (95% athlete-cluster CI 0.5204–0.8848) and PR-AUC 0.0210 (95% CI 0.0040–0.0449; no-skill reference 0.0067), with Brier score 0.2038. An otherwise identical unweighted model had similar ranking discrimination (ROC-AUC 0.7550; PR-AUC 0.0199) but Brier score 0.0080, showing that class weighting materially affected the score scale. At the training-derived top-5% operating point, 1508/17,535 player-days (8.60%) were flagged; precision was 0.0252 and 9/12 episodes were detected (75%; exact 95% CI 42.8–94.5%), with 1470 false-alert player-days, 640 false-alert episodes, and 565/2550 observed athlete-weeks (22.16%) containing at least one false alert. In the personalized task, hold-out macro F1 decreased from 0.9773 to 0.3064 after simultaneous removal of four strongly target-associated predictors; in the biomechanical task, no non-dummy model exceeded the stratified dummy classifier in repeated cross-validation.
Conclusions
SoccerMon showed a possible ranking signal, but its magnitude was imprecisely estimated and class-weighted outputs were not valid absolute-risk estimates without calibration. Internal season-to-season validation in the same athlete panel, only 12 test episodes, temporal dataset shift, and substantial false-alert burden preclude claims of transportability or clinical utility. The supporting supplied-label cases illustrate contrasting proxy dependence and absence of reproducible signal beyond a non-informative baseline. The study is a methodological evaluation rather than a deployment-ready injury-forecasting system.