构建首个面向多模态交互的人类主体伪造视频基准,支持可解释的伪造推理。
HumanForge: A Human-Centric Deepfake Video Benchmark with Multi-Agent Forgery Rationales

- 设计多智能体协作流水线,自动生成并标注1.8万条伪造视频
- 涵盖音频、姿态、语义和交互四类场景,支持跨生成器泛化挑战测试
- 提供自然语言对比性伪造理由,适合研究可解释视频取证的学者
视频扩散模型与时序编辑工具的快速发展催生了高度逼真的人类主体视频生成,对数字内容取证构成前所未有的挑战。现有基准主要聚焦于换脸或全局文本到视频合成,忽视了多模态对齐及复杂的人-物、人-人交互等关键维度。为此,我们提出HumanForge,一个统一的、大规模的、多范式的人类主体伪造视频基准,包含超过18,000条合成视频,覆盖四种不同场景:音频驱动、姿态驱动、语义驱动和交互。为避免人工标注或盲目的单一提示,我们提出Gen2Anno(生成到标注)框架,协调六个专业化智能体——从驱动资产分析到MoE参考分析、闭环验证——动态执行视频合成并生成结构化标注,包括二元真实性标签、生成模型归属和自然语言对比性伪造理由。通过系统性对比生成来源预期状态与实际视觉观察,该框架生成逻辑自洽的取证推理链。使用先进传统检测器和视觉-语言模型的广泛基准测试表明,HumanForge在跨生成器泛化、扰动鲁棒性和可解释推理方面带来显著挑战。代码与数据集将公开发布。
原文摘要 · Abstract (English)
Rapid advancements in video diffusion models and temporal editing tools have enabled the generation of highly realistic human-centric videos, presenting unprecedented challenges to digital content forensics. Existing benchmarks primarily focus on face-swapping or global text-to-video synthesis, overlooking the crucial dimensions of multimodal alignment and complex human-object or human-human interactions. To address these limitations, we introduce HumanForge, a unified, large-scale, and multi-paradigm human-centric video forgery benchmark containing over 18,000 synthesized videos across four distinct scenarios: audio-driven, pose-driven, semantic-driven, and interaction. To construct and annotate this dataset without labor-intensive manual labeling or blind monolithic prompting, we propose Gen2Anno (Generation-to-Annotation), a cooperative multi-agent pipeline. Gen2Anno orchestrates six specialized agents-ranging from driving asset profiling to MoE-based reference analysis and closed-loop verification-to dynamically execute video synthesis and produce structured annotations containing binary authenticity labels, generative model attribution, and natural-language contrastive forgery rationales. By systematically contrasting expected states derived from generation provenance with actual visual observations, the framework generates logically grounded forensic reasoning chains. Extensive benchmarks using state-of-the-art traditional detectors and Vision-Language Models demonstrate the significant challenges of cross-generator generalization, perturbation robustness, and explainable reasoning on HumanForge. The code and dataset will be publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。