用语义反馈提升视频推理的精准定位能力
SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards

- 引入语义证据奖励机制,通过视觉语言模型验证证据相关性与定位质量
- 在V-STAR上达到49.6% mLGM,比基线提升3.0点
- 无需密集框标注,可直接在标准视频问答数据上训练
视频多模态大模型在细粒度时空推理中常因依赖无关帧或物体而生成错误答案。尽管输出时空证据是可行方向,但现有强化学习框架多依赖仅基于几何(IoU)的奖励,易受边界扰动影响且忽略语义对齐。为此,我们提出语义证据奖励(SER),将时空证据定位重构为约束验证任务。SER不计算像素级重叠,而是利用裁判型视觉语言模型从相关性和定位质量两个维度评估模型生成的证据主张,并结合时间惩罚。该设计降低对密集框标注的依赖,支持直接在标准视频问答数据上训练。在V-STAR基准上,SER实现49.6% mLGM,较强基线Open-o3-Video提升3.0点,验证了其在提升答案准确率与证据定位能力方面的潜力。
原文摘要 · Abstract (English)
Video MLLMs often struggle with fine-grained spatio-temporal reasoning, sometimes generating correct answers based on irrelevant frames or objects. Although outputting spatio-temporal evidence during reasoning is a promising direction, existing RL frameworks typically rely on geometry-only (IoU) rewards, which can be sensitive to boundary perturbations and overlook semantic alignment. To address this, we propose Semantic Evidence Reward (SER), which reformulates spatio-temporal evidence grounding as a constrained verification task. Instead of computing pixel-level overlap, SER uses a referee VLM as a local checker to evaluate model-generated evidence claims across two dimensions: relevance and localization quality, combined with a temporal penalty. This design reduces the reliance on dense box annotations and enables training directly on standard video QA data. On the V-STAR benchmark, SER achieves 49.6% mLGM, improving by 3.0 points over the strong evidence-grounded baseline Open-o3-Video, demonstrating its potential in enhancing both answer accuracy and evidence grounding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。