用强化学习让AI看图识破全身合成视频,无需大量标注数据。
AvatarShield: Visual Reinforcement Learning for Human-Centric Synthetic Video Detection
- 用二元标签训练大模型,通过分组相对策略优化提升推理能力。
- 在15000条视频上测试,对九种主流生成方法的检测准确率超现有技术。
- 适合关注视频真伪验证与AI可解释性的研究者和安全团队。
人工智能生成内容的进展使得高度逼真的合成视频日益普遍,尤其在涉及语音、手势和全身动作的人类中心场景中,严重威胁信息真实性与公众信任。与仅关注局部面部篡改的DeepFake不同,当前的人类中心视频生成技术可合成完整人体并控制其运动,实现与环境、物体甚至他人的复杂交互。然而,现有检测方法大多忽视此类全身合成内容带来的风险。尽管有研究尝试利用大语言模型(LLM)进行可解释性检测,但依赖监督微调导致标注偏差、幻觉监督和泛化能力下降。为此,我们提出AvatarShield,一种无需密集文本标注的多模态人类中心合成视频检测框架,采用分组相对策略优化(Group Relative Policy Optimization),使LLM仅凭二元标签即可发展出强推理能力。该框架结合离散视觉塔用于高层次语义不一致检测,以及残差提取器进行细粒度伪影分析。我们还构建了FakeHumanVid,一个包含15,000条真实与合成视频的大规模基准数据集,涵盖九种由文本、姿态或音频驱动的先进生成方法。大量实验表明,AvatarShield在域内与跨域设置下均优于现有方法。
原文摘要 · Abstract (English)
Recent advances in Artificial Intelligence Generated Content have led to highly realistic synthetic videos, particularly in human-centric scenarios involving speech, gestures, and full-body motion, posing serious threats to information authenticity and public trust. Unlike DeepFake techniques that focus on localized facial manipulation, human-centric video generation methods can synthesize entire human bodies with controllable movements, enabling complex interactions with environments, objects, and even other people. However, existing detection methods largely overlook the growing risks posed by such full-body synthetic content. Meanwhile, a growing body of research has explored leveraging LLMs for interpretable fake detection, aiming to explain decisions in natural language. Yet these approaches heavily depend on supervised fine-tuning, which introduces limitations such as annotation bias, hallucinated supervision, and weakened generalization. To address these challenges, we propose AvatarShield, a novel multimodal human-centric synthetic video detection framework that eliminates the need for dense textual supervision by adopting Group Relative Policy Optimization, enabling LLMs to develop reasoning capabilities from simple binary labels. Our architecture combines a discrete vision tower for high-level semantic inconsistencies and a residual extractor for fine-grained artifact analysis. We further introduce FakeHumanVid, a large-scale benchmark containing 15K real and synthetic videos across nine state-of-the-art human generation methods driven by text, pose, or audio. Extensive experiments demonstrate that AvatarShield outperforms existing methods in both in-domain and cross-domain settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。