构建首个需联网检索的视频假信息检测基准,挑战模型跨视频验证能力。
When Seeing Is Not Believing -- A Benchmark for Search-Grounded Video Misinformation Detection

- 通过联网搜索相关视频,对比识别篡改内容,突破仅看视频无法发现假信息的局限
- 9个前沿多模态模型在该基准上点级准确率仅61.43%,视频级准确率43.24%
- 适合研究虚假信息检测、多模态推理与搜索增强系统的学者使用
视频假信息正从表面伪造转向语义与证据层面的操纵:真实画面可能被选择性剪辑、时间顺序错乱、跨源拼接或叠加生成内容以构建虚假叙事。这类依赖证据的篡改无法仅通过观看视频本身识别,因缺失、错位、替换或再上下文化的证据位于视频之外。我们提出 extbf{EVID-Bench},一个面向搜索增强型视频假信息检测的基准,要求系统在开放网络中检索相关视频,并通过跨视频比对识别错误信息。EVID-Bench 包含 222 个视频,涵盖 9 种篡改类型,分布在人工智能生成、单源编辑和多源编辑三类中。所有样本均经人工验证,无法被前沿模型通过视觉检查识别。我们评估了九个前沿多模态模型,采用检索增强验证基线。最佳系统点级准确率为 61.43%,视频级准确率为 43.24%;其中人工智能生成类篡改尤其难辨。错误分析显示,模型常聚焦无关锚点、误将合成内容归为剪辑、过早终止搜索而未能完整解释篡改逻辑。
原文摘要 · Abstract (English)
Video misinformation increasingly operates at the semantic and evidential level: authentic footage may be selectively edited, temporally reordered, spliced across sources, or augmented with AI-generated content to construct false narratives. Such evidence-dependent manipulations cannot be reliably verified from the input video alone, because the missing, reordered, replaced, or recontextualized evidence lies outside the video itself. We introduce \textbf{EVID-Bench}, a benchmark for search-grounded video misinformation detection, where a system must search the open web for related videos and identify what information is false through cross-video comparison. EVID-Bench comprises 222 videos spanning 9 manipulation types across 3 categories: AI generation, single-source editing, and multi-source editing. All samples are verified to be undetectable by frontier models through visual inspection alone. We evaluate nine frontier multimodal models using a retrieval-augmented verification baseline. The best system achieves only 61.43\% point-level accuracy and 43.24\% video-level accuracy, while AI-generated manipulations remain especially challenging. Error analysis reveals recurring challenges: models fixate on irrelevant anchors, misattribute synthetic content to editorial splicing, and terminate search prematurely before fully explaining the manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。