用结构化奖励提升视频多模态模型的视觉与时间一致性
Reinforcing Consistency in Video MLLMs with Structured Rewards
- 设计基于事实与时间单元的分层奖励机制
- 在多个基准上显著减少幻觉,提升事实支持度
- 适合关注视频理解真实性与对齐的研究者
多模态大语言模型在视频理解任务中取得显著进展,但其看似合理的输出常缺乏可靠的视觉与时间依据:模型可能虚构物体存在、误标属性或忽略重复事件,却仍生成整体通顺的描述。本文通过成分一致性审计,将句子分解为支持性事实与时间命题,发现即使高层关系正确,底层属性和存在性支持仍常缺失。这表明标准句级监督对忠实视频理解是弱代理。进一步使用强化学习时,句级奖励过于粗略,难以定位具体对齐失败。为此,我们引入由事实与时间单元构成的结构化奖励:(1)实例感知场景图奖励用于物体、属性与关系;(2)时间奖励用于事件顺序与重复;(3)视频接地VQA奖励用于层级自验证。在时间理解、通用视频理解及抗幻觉基准上,该目标在开源模型上均实现一致提升,表明结构化奖励塑造是实现更忠实视频理解的有效路径。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have achieved remarkable progress in video understanding. However, seemingly plausible outputs often suffer from poor visual and temporal grounding: a model may fabricate object existence, assign incorrect attributes, or collapse repeated events while still producing a globally reasonable caption or answer. We study this failure mode through a compositional consistency audit that decomposes a caption into supporting factual and temporal claims, investigating whether a correct high-level prediction is actually backed by valid lower-level evidence. Our top-down audit reveals that even correct root relational claims often lack reliable attribute and existence support. This indicates that standard sentence-level supervision is a weak proxy for faithful video understanding. Furthermore, when turning to reinforcement learning (RL) for better alignment, standard sentence-level rewards often prove too coarse to accurately localize specific grounding failures. To address this, we replace generic sentence-level rewards with a structured reward built from factual and temporal units. Our training objective integrates three complementary components: (1) an instance-aware scene-graph reward for factual objects, attributes, and relations; (2) a temporal reward for event ordering and repetition; and (3) a video-grounded VQA reward for hierarchical self-verification. Across temporal, general video understanding, and hallucination-oriented benchmarks, this objective yields consistent gains on open-source backbones. These results suggest that structured reward shaping is a practical route to more faithful video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。