聚焦接触瞬间,提升单视角假摔识别准确率。
CAS-FD: Contact-Aware Temporal Sampling for Single-View Foul vs Dive Recognition

- 关注物理接触时刻,动态采样关键帧
- 达到86.0%准确率,较传统方法提升12个百分点
- 适合体育裁判辅助与视频分析研究者
在缺乏多视角摄像机的单视角直播画面中,区分足球比赛中的真实犯规与假摔仍是极具挑战性的细粒度识别任务。本文构建了一个包含600个片段的平衡数据集,并提出一种接触感知的时间采样策略,使模型注意力集中在身体接触瞬间,而非均匀处理所有帧。该方法在测试集上取得86.0%的准确率和0.860的宏平均F1值,较无接触感知的基线方法提升12个百分点,且在未见数据上优势更明显。通过与人工标注对比评估各模块性能,明确了系统的优势与局限。研究成果包括可复现的单视角识别流程、公开的数据集与评估框架,为广播足球视频中的接触事件识别提供坚实基础。代码与数据已开源。
原文摘要 · Abstract (English)
Distinguishing a genuine foul from a simulated dive in football remains one of the sport's most contested fine-grained recognition problems, especially when such decisions have to be from a single broadcast view without multi-view camera angle. We introduce a balanced 600-clip single-view Foul/Dive dataset and show that contact-aware sampling concentrating the model's attention around the moment of physical contact rather than treating all frames equally yields substantially improved recognition of this contact- specific problem. The proposed approach achieves 86.0% accuracy and macro-F1 0.860 on the held-out test split, a 12 percentage- point gain over contact-unaware alternatives that grows further on unseen data. We also evaluate each pipeline component against human annotations, establishing where and why the system suc- ceeds and fails. The result is a documented dataset, a reproducible single-view pipeline, and a grounded evaluation framework for fine-grained contact-event recognition in broadcast football footage. The dataset and code are available at https://github.com/hossain- tamim/contact-aware-dive.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。