用多模态嵌入区分真实与AI生成的人体动作,抗压缩和分辨率篡改。
Human Action CLIPs: Detecting AI-generated Human Motion
- 基于多模态语义嵌入检测人体动作真伪
- 在7个文本生成视频模型上准确率超90%
- 适合内容安全、防伪造领域研究者使用
AI生成视频正逐步逼近真实感,为防范其恶意应用,本文提出一种基于多模态语义嵌入的鲁棒方法,用于区分真实与AI生成的人体动作。该方法对分辨率调整、压缩等常见篡改手段具有强抗性。我们在DeepAction数据集上进行评估,该数据集包含由7个文本到视频模型生成的真人动作视频及对应真实视频,数据集可通过Hugging Face获取(https://www.huggingface.co/datasets/faridlab/deepaction_v1),支持学术研究。
原文摘要 · Abstract (English)
AI-generated video generation continues its journey through the uncanny valley to produce content that is increasingly perceptually indistinguishable from reality. To better protect individuals, organizations, and societies from its malicious applications, we describe an effective and robust technique for distinguishing real from AI-generated human motion using multi-modal semantic embeddings. Our method is robust to the types of laundering that typically confound more low- to mid-level approaches, including resolution and compression attacks. This method is evaluated against DeepAction, a custom-built, open-sourced dataset of video clips with human actions generated by seven text-to-video AI models and matching real footage. The dataset is available under an academic license at https://www.huggingface.co/datasets/faridlab/deepaction_v1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。