用固定时间点构建球员单次训练/比赛的表征,避免错误的分钟级伤情标注。
Landmark-Based Discrimination of Injury-Associated Athlete-Sessions from Minute-Resolution Multimodal Football Monitoring Data

- 以固定时间点为基准,整合各时段数据生成球员-会话统一表征
- 模型在不同时间点上对伤情的判别能力有限,AUC最高仅0.607
- 适用于高精度运动监控中缺乏精确伤情发生时间的场景
运动员监测数据通常按分钟记录,但伤病信息仅提供整个会话是否关联伤病。这导致建模难题:若将会话标签分配给每分钟数据,会错误暗示每个时刻都已知伤病状态,而实际伤情发作时间未知。本文提出一种基于固定时间点(如10、20、30分钟)的单一表征方法,每名运动员-会话仅生成一个表征,利用该时间点前的所有信息进行判断。此方法保持标签在会话层面,避免了无依据的分钟级监督。基于2020年SoccerMon数据,分析48名精英女子足球运动员的3,743个球员-会话,其中22个为伤情相关会话,涉及5名运动员。通过球员不重叠验证、运动员聚类自助法不确定性评估、共用队列敏感性分析、替代负样本分组、等权重设置,以及逻辑回归、随机森林和XGBoost基准对比,结果显示主模型(CUM+DYN逻辑回归)在不同时间点的ROC-AUC为0.367–0.607,PR-AUC为0.0080–0.0150,不确定性较大;包含预会话信息的表征在多个时间点有更高估计值,但仍不确定。
原文摘要 · Abstract (English)
Athlete monitoring data may be recorded minute by minute throughout a match or training session, while injury information may only indicate whether the entire session was injury-associated. This creates a modelling problem: assigning the same session-level label to every minute would imply that injury status is known at each exact time, even though within-session injury onset is unknown. Our novelty is a fixed-landmark, one-representation-per-athlete-session formulation that directly addresses this mismatch. Instead of labelling every minute, we construct one representation per athlete-session at each landmark using information observed up to that point. This keeps the target at the session level and avoids unsupported minute-level injury supervision. A landmark is a fixed time point within the same session, such as 10, 20, or 30 minutes. At each landmark, we assess whether the whole session is injury-associated or non-injury-associated and examine how discrimination changes as more within-session information becomes available. Using 2020 SoccerMon data, we analyse 3,743 athlete-sessions from 48 elite women's football athletes, including 22 injury-associated sessions from five athletes. We evaluate pre-session, cumulative, dynamic, and combined representations with athlete-disjoint validation, athlete-cluster bootstrap uncertainty, common-cohort sensitivity analysis, alternative negative-athlete fold allocations, equal-athlete weighting, and Logistic Regression, Random Forest, and XGBoost benchmarks. Primary CUM+DYN Logistic Regression yields ROC-AUC 0.367-0.607 and PR-AUC 0.0080-0.0150 across landmarks, with wide uncertainty. PRE-containing representations show higher point estimates at several landmarks but remain uncertain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。