arXiv:2506.00868cs.MMcs.CV2025-06被引 9

构建首个面向人物中心的高阶视觉与概念篡改数据集,揭示当前检测模型盲区

Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations

  • 基于视觉语言模型生成84万余张语义化篡改图像,聚焦人物与场景意义层面修改
  • 现有顶尖检测模型与人类观察者在识别此类微妙篡改时准确率均不足60%
  • 适合研究深层伪造检测、可信生成内容评估及人机认知差异的学者使用

近年来生成式人工智能技术的快速发展推动了高度逼真的深度伪造内容生成。尽管已有诸多努力,研究社区仍缺乏一个大规模且具备推理能力的深度伪造基准数据集,专门用于人物中心的对象、上下文和场景篡改研究。本文提出MultiFakeVerse,一个大规模的人物中心深度伪造数据集,包含845,286张通过视觉语言模型(VLM)生成的篡改图像,其指令源于对个体或场景上下文元素的修改建议,这些元素会影响人类对重要性、意图或叙事的理解。该方法实现语义化、上下文感知的修改,如改变动作、场景和人-物交互,而非传统低层次的身份替换或区域编辑。实验表明,当前最先进的深度伪造检测模型和人类观察者难以识别这类细微但具有意义的篡改。代码与数据集已开源于GitHub。

原文摘要 · Abstract (English)

The rapid advancement of GenAI technology over the past few years has significantly contributed towards highly realistic deepfake content generation. Despite ongoing efforts, the research community still lacks a large-scale and reasoning capability driven deepfake benchmark dataset specifically tailored for person-centric object, context and scene manipulations. In this paper, we address this gap by introducing MultiFakeVerse, a large scale person-centric deepfake dataset, comprising 845,286 images generated through manipulation suggestions and image manipulations both derived from vision-language models (VLM). The VLM instructions were specifically targeted towards modifications to individuals or contextual elements of a scene that influence human perception of importance, intent, or narrative. This VLM-driven approach enables semantic, context-aware alterations such as modifying actions, scenes, and human-object interactions rather than synthetic or low-level identity swaps and region-specific edits that are common in existing datasets. Our experiments reveal that current state-of-the-art deepfake detection models and human observers struggle to detect these subtle yet meaningful manipulations. The code and dataset are available on \href{https://github.com/Parul-Gupta/MultiFakeVerse}{GitHub}.

深度伪造数据集视觉语言模型生成内容

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。