测试视频大模型在诱导干扰下的真实性和事实性可靠性,发现高准确率不等于强鲁棒性。
INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs
- 构建9800个问答样本,分维度评估视频模型的忠实度与事实性。
- 证据篡改和时间干预使模型性能大幅下降,部分开源模型对时序问题近乎无敏感性。
- 适用于评估视频大模型在真实复杂场景中的可信度,尤其关注时序敏感任务。
尽管进展迅速,视频大语言模型仍因幻觉问题而不可靠,即输出内容与视频证据矛盾(忠实性)或违背可验证的世界知识(事实性)。现有基准对事实性幻觉覆盖有限,且主要在干净环境下评估。我们提出 extsc{INFACT},一个包含9,800个问答实例的诊断基准,具有细粒度的忠实性与事实性分类体系,涵盖真实与合成视频。该基准在四种模式下评估模型:基础(干净)、视觉退化、证据篡改和时序干预(针对时序敏感项)。通过抗扰率(RR)和时序敏感度得分(TSS)量化模型在诱导模式下的可靠性。对14个代表性视频大模型的实验表明,基础模式下的高准确率并不能可靠转化为诱导模式下的高可靠性,证据篡改降低稳定性,时序干预导致最大性能下降。值得注意的是,许多开源基线在事实性任务上的时序敏感度接近零,显示其在时序敏感问题上存在显著时序惯性。
原文摘要 · Abstract (English)
Despite rapid progress, Video Large Language Models (Video-LLMs) remain unreliable due to hallucinations, which are outputs that contradict either video evidence (faithfulness) or verifiable world knowledge (factuality). Existing benchmarks provide limited coverage of factuality hallucinations and predominantly evaluate models only in clean settings. We introduce \textsc{INFACT}, a diagnostic benchmark comprising 9{,}800 QA instances with fine-grained taxonomies for faithfulness and factuality, spanning real and synthetic videos. \textsc{INFACT} evaluates models in four modes: Base (clean), Visual Degradation, Evidence Corruption, and Temporal Intervention for order-sensitive items. Reliability under induced modes is quantified using Resist Rate (RR) and Temporal Sensitivity Score (TSS). Experiments on 14 representative Video-LLMs reveal that higher Base-mode accuracy does not reliably translate to higher reliability in the induced modes, with evidence corruption reducing stability and temporal intervention yielding the largest degradation. Notably, many open-source baselines exhibit near-zero TSS on factuality, indicating pronounced temporal inertia on order-sensitive questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。