arXiv:2507.06523cs.CVcs.CL2025-07被引 2

提出统一评估视频生成与理解中事实一致性的框架FIFA。

FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation

  • 构建时空语义依赖图,提取并验证生成内容中的事实
  • FIFA评估结果更贴近人类判断,显著降低幻觉率
  • 配套后修正工具可有效提升文本与视频生成的准确性

视频多模态大模型在视频到文本和文本到视频任务上取得了显著进展,但常出现与视觉输入矛盾的幻觉。现有评估方法仅针对单一任务(如视频到文本),且无法评估开放式自由生成中的幻觉。为此,我们提出FIFA——一个统一的事实性评估框架,通过提取全面描述性事实,利用时空语义依赖图建模其语义关联,并借助视频问答模型进行验证。我们还引入后修正(Post-Correction)工具框架,对幻觉内容进行基于工具的修正。大量实验表明,FIFA比现有方法更接近人类判断,且后修正能有效提升文本与视频生成的事实一致性。

原文摘要 · Abstract (English)

Video Multimodal Large Language Models (VideoMLLMs) have achieved remarkable progress in both Video-to-Text and Text-to-Video tasks. However, they often suffer fro hallucinations, generating content that contradicts the visual input. Existing evaluation methods are limited to one task (e.g., V2T) and also fail to assess hallucinations in open-ended, free-form responses. To address this gap, we propose FIFA, a unified FaIthFulness evAluation framework that extracts comprehensive descriptive facts, models their semantic dependencies via a Spatio-Temporal Semantic Dependency Graph, and verifies them using VideoQA models. We further introduce Post-Correction, a tool-based correction framework that revises hallucinated content. Extensive experiments demonstrate that FIFA aligns more closely with human judgment than existing evaluation methods, and that Post-Correction effectively improves factual consistency in both text and video generation.

视频生成事实一致性评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。