arXiv:2604.25584cs.AI2026-04

提出双层多模态事实验证框架,提升视频描述的准确性评估。

DualFact+: A Multimodal Fact Verification Framework for Procedural Video Understanding

论文配图:DualFact+: A Multimodal Fact Verification Framework for Procedural Video Understanding
图 1 · 摘自论文原文
  • 分概念与上下文两类事实,分别评估语义角色与视觉实化。
  • 在YouCook3-Fact和CraftBench-Fact上发现模型常遗漏关键信息。
  • 结合视觉证据更贴近人类判断,适合评估多模态生成真实性。

我们提出DualFact,一种双层多模态事实性评估框架,用于程序性视频字幕生成。该框架将事实正确性分为概念事实(如动作、原料、工具、位置等抽象语义角色)与上下文事实(这些角色在视频中的具体谓词-论元实现)。为实现完整且角色一致的评估,引入隐式论元增强(VIA)与对比事实集。双重实例化模式分别为:基于文本证据的DualFact-T,以及基于视频视觉证据的DualFact-V。在YouCook3-Fact与CraftBench-Fact上的实验表明,当前顶尖多模态语言模型虽生成流畅字幕,但常存在系统性遗漏与角色不一致问题。DualFact与人工事实判断相关性更高,尤其在上下文事实方面;相比仅依赖字幕的评估,基于视频的验证更真实反映幻觉程度。整体而言,DualFact提供了一种可解释且与人类对齐的评估协议,揭示了多模态事实锚定中的持续挑战,超越表面流畅性。

原文摘要 · Abstract (English)

We introduce DualFact, a dual-layer, multimodal factuality evaluation framework for procedural video captioning. DualFact separates factual correctness into conceptual facts, capturing abstract semantic roles (e.g., Action, Ingredient, Tool, Location), and contextual facts, capturing their grounded predicate-argument realizations in video. To support complete and role-consistent evaluation, DualFact incorporates implicit argument augmentation (VIA) and contrastive fact sets. We instantiate DualFact in two modes: DualFact-T, which verifies facts against textual evidence, and DualFact-V, which verifies facts against video-grounded visual evidence. Experiments on YouCook3-Fact and CraftBench-Fact show that state-of-the-art multimodal language models produce fluent but often factually incomplete captions, with systematic omissions and role-level inconsistencies. DualFact correlates more strongly with human factuality judgments than standard metrics, particularly for contextual facts, and reveals that caption-only evaluation overestimates hallucinations compared to video-grounded verification. Overall, DualFact offers an interpretable and human-aligned evaluation protocol that highlights persistent challenges in multimodal factual grounding, extending beyond surface-level fluency.

多模态事实验证视频理解生成评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。