arXiv:2512.00336cs.CV2025-12被引 1

首个面向多模态伪造音视频的检测基准数据集

MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection

  • 构建三类真实伪造模式的音视频合成数据
  • 覆盖真人、动物、物体等四类内容,支持四种模态组合
  • 适配扩散模型生成的高保真音视频,助力可信检测

AI生成的多模态音视频内容快速发展,引发信息真实性与安全担忧。现有合成视频数据集多仅关注视觉模态,少数含音频的数据也局限于人脸深度伪造,难以应对通用多模态伪造场景,严重制约可信检测系统的发展。为此,我们提出首个专为检测多模态AI生成音视频设计的基准数据集——MVAD。该数据集具备三大特征:(1) 真实多模态性,基于三种现实音视频伪造模式生成;(2) 高感知质量,采用多样化的先进生成模型实现;(3) 全面多样性,涵盖真实与动漫风格,四类内容(人物、动物、物体、场景)及四种视频-音频多模态数据类型。数据集将开源发布于 https://github.com/HuMengXue0104/MVAD。

原文摘要 · Abstract (English)

The rapid advancement of AI-generated multimodal video-audio content has raised significant concerns regarding information security and content authenticity. Existing synthetic video datasets predominantly focus on the visual modality alone, while the few incorporating audio are largely confined to facial deepfakes--a limitation that fails to address the expanding landscape of general multimodal AI-generated content and substantially impedes the development of trustworthy detection systems. To bridge this critical gap, we introduce the Multimodal Video-Audio Dataset (MVAD), the first comprehensive dataset specifically designed for detecting AI-generated multimodal video-audio content. Our dataset exhibits three key characteristics: (1) genuine multimodality with samples generated according to three realistic video-audio forgery patterns; (2) high perceptual quality achieved through diverse state-of-the-art generative models; and (3) comprehensive diversity spanning realistic and anime visual styles, four content categories (humans, animals, objects, and scenes), and four video-audio multimodal data types. Our dataset will be available at https://github.com/HuMengXue0104/MVAD.

音视频伪造多模态检测数据集AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。