通过任务间注意力一致性提升体育教学视频的时间定位精度
Learning Consistent Temporal Grounding between Related Tasks in Sports Coaching
- 利用相关任务注意力图的一致性约束,无需额外标注
- 在三项体育任务中准确率提升最高达14.1%
- 适合需要精准时间定位的视频理解场景
视频大模型常关注无关帧,这对需要精确时间定位的体育教学任务尤为不利。然而获取帧级标注困难:人工标注成本高,模型自动生成不可靠。本文提出利用相关任务(如生成与验证)必须关注相同帧的特性,通过选择性视觉注意力图施加自一致性目标来改善时间定位,无需额外标注。基于VidDiffBench(含真实关键帧标注)验证了注意力错位是主要瓶颈。实验显示,采用该目标训练,在Exact、FitnessQA和ExpertAF三个体育教学任务上,相比监督微调分别获得+3.0%、+14.1%准确率提升及+0.9 BERTScore,甚至超越闭源模型。
原文摘要 · Abstract (English)
Video-LLMs often attend to irrelevant frames, which is especially detrimental for sports coaching tasks requiring precise temporal grounding. Yet obtaining frame-level supervision is challenging: expensive to collect from humans and unreliable from other models. We improve temporal grounding without additional annotations by exploiting the observation that related tasks, such as generation and verification, must attend to the same frames. We enforce this via a self-consistency objective over select visual attention maps of tightly-related tasks. Using VidDiffBench, which provides ground-truth keyframe annotations, we first validate that attention misallocation is a significant bottleneck. We then show that training with our objective yields gains of +3.0%, +14.1% accuracy and +0.9 BERTScore over supervised finetuning across three sports coaching tasks: Exact, FitnessQA, and ExpertAF, even surpassing closed-source models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。