arXiv:2511.21251cs.CV2025-11被引 1

首个覆盖多类型音视频伪造的综合评测基准,助力AI识别真实世界伪造场景。

AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMs

  • 构建多阶段混合伪造框架生成高质量多样化音视频伪造数据。
  • 包含12,000个问题、7类伪造类型与4级标注,支持多任务评估。
  • 首次系统评测11个音视频大模型在伪造检测中的表现与短板。

音视频伪造威胁正从以人类为中心的深度伪造,扩展到复杂自然场景中更多样化的篡改形式。然而,现有评测基准仍局限于基于DeepFake的伪造和单一粒度标注,难以捕捉现实伪造场景的多样性与复杂性。为此,我们提出AVFakeBench,首个涵盖人类与非人类主体的综合性音视频伪造检测基准。该基准包含12,000个精心设计的音视频问答,覆盖七种伪造类型与四个层级的标注。为确保伪造内容的高质量与多样性,我们提出一种多阶段混合伪造框架,结合专有任务规划模型与专家生成模型实现精准操控。基准建立涵盖二分类判断、伪造类型分类、伪造细节定位与解释性推理的多任务评估体系。我们在该基准上评估了11个音视频大语言模型(AV-LMMs)及2种主流检测方法,验证了AV-LMMs作为新兴伪造检测器的潜力,同时揭示其在细粒度感知与推理能力上的显著不足。

原文摘要 · Abstract (English)

The threat of Audio-Video (AV) forgery is rapidly evolving beyond human-centric deepfakes to include more diverse manipulations across complex natural scenes. However, existing benchmarks are still confined to DeepFake-based forgeries and single-granularity annotations, thus failing to capture the diversity and complexity of real-world forgery scenarios. To address this, we introduce AVFakeBench, the first comprehensive audio-video forgery detection benchmark that spans rich forgery semantics across both human subject and general subject. AVFakeBench comprises 12K carefully curated audio-video questions, covering seven forgery types and four levels of annotations. To ensure high-quality and diverse forgeries, we propose a multi-stage hybrid forgery framework that integrates proprietary models for task planning with expert generative models for precise manipulation. The benchmark establishes a multi-task evaluation framework covering binary judgment, forgery types classification, forgery detail selection, and explanatory reasoning. We evaluate 11 Audio-Video Large Language Models (AV-LMMs) and 2 prevalent detection methods on AVFakeBench, demonstrating the potential of AV-LMMs as emerging forgery detectors while revealing their notable weaknesses in fine-grained perception and reasoning.

伪造检测音视频大模型评测多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。