arXiv:2511.18463cs.CV2025-11

分离视觉感知与推理,提升视频理解的准确性与可验证性。

Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding

  • 将感知信息结构化为带时间戳的描述单元,独立于推理过程。
  • 引入事实感知评估器,使幻觉检测性能媲美GPT-4o。
  • 适用于需要高可靠性的视频理解任务,如医疗或自动驾驶场景。

视频大模型通过生成中间推理文本来提升复杂视频的理解能力,但可靠的推理依赖于准确的视频感知。现有方法中,感知证据与推理文本紧密耦合,难以对感知过程进行直接监督。本文提出解耦感知与逻辑(DPL)框架,将感知表示为包含时间戳和视觉描述的固定格式证据单元,实现感知内容的直接提取,并简化视频片段与奖励评估之间的对齐。基于DPL,设计感知奖励机制,同时鼓励抗幻觉能力和基于感知的推理。采用事实感知评估器(FAE)提供反幻觉评分,其幻觉检测性能接近GPT-4o。此外,通过将感知结果与问题输入参考模型,验证推理一致性。实验表明,视频-DPL在3B和7B规模下均显著提升微调后性能,并具备更高的数据效率。

原文摘要 · Abstract (English)

Video Large Language Models improve reasoning over complex videos by generating intermediate reasoning text. However, reliable reasoning depends on accurate video perception. In existing approaches, perception evidence is intertwined with reasoning text, making it difficult to directly supervise the perception process. We argue that reliable supervision requires explicitly separating perception evidence from reasoning so that perception can be verified independently. To supervise perception directly, we propose Decoupled Perception and Logic (DPL), which represents perception as fixed-format evidence units containing timestamps and visual descriptions. This structured representation enables direct extraction of perception content and simplifies alignment between video segments and reward evaluation. Building on DPL, we introduce a perception reward that encourages both hallucination resistance and perception-based reasoning. An Factual-Aware Evaluator (FAE) provides anti-hallucination scores and achieves hallucination evaluation performance comparable to GPT-4o. In addition, we validate reasoning consistency by feeding perception results and questions into a reference model. Experiments show that, by providing reliable process rewards, Video-DPL consistently improves post-training performance at both 3B and 7B scales, while delivering higher data efficiency.

视频理解幻觉抑制大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。