arXiv:2411.09266cs.CVcs.AI2024-11被引 11

用ChatGPT检测音视频伪造,表现媲美专业模型。

How Good is ChatGPT at Audiovisual Deepfake Detection: A Comparative Study of ChatGPT, AI Models and Human Perception

  • 通过提示工程让ChatGPT分析音视频中的时空不一致特征。
  • 在基准数据集上检测准确率接近先进多模态模型。
  • 适合需要可解释性、低成本的伪造检测场景。

音视频深度伪造因视觉与听觉篡改难以被肉眼察觉,传统单模态检测方法效果有限。尽管多模态取证模型能力更强,但需大量训练数据,计算成本高,且缺乏可解释性,泛化能力弱。本研究评估大语言模型(如ChatGPT)在识别音视频伪造内容中的表现,考察其对视觉与听觉伪影的感知能力。在标准多模态深度伪造数据集上开展大量实验,对比ChatGPT、前沿多模态模型及人类感知的检测性能。结果表明,领域知识与提示工程对基于LLM的伪造检测至关重要。与端到端学习方法不同,ChatGPT能识别跨模态或模态内的空间与时空不一致。同时讨论了其在多媒体取证任务中的局限性。

原文摘要 · Abstract (English)

Multimodal deepfakes involving audiovisual manipulations are a growing threat because they are difficult to detect with the naked eye or using unimodal deep learningbased forgery detection methods. Audiovisual forensic models, while more capable than unimodal models, require large training datasets and are computationally expensive for training and inference. Furthermore, these models lack interpretability and often do not generalize well to unseen manipulations. In this study, we examine the detection capabilities of a large language model (LLM) (i.e., ChatGPT) to identify and account for any possible visual and auditory artifacts and manipulations in audiovisual deepfake content. Extensive experiments are conducted on videos from a benchmark multimodal deepfake dataset to evaluate the detection performance of ChatGPT and compare it with the detection capabilities of state-of-the-art multimodal forensic models and humans. Experimental results demonstrate the importance of domain knowledge and prompt engineering for video forgery detection tasks using LLMs. Unlike approaches based on end-to-end learning, ChatGPT can account for spatial and spatiotemporal artifacts and inconsistencies that may exist within or across modalities. Additionally, we discuss the limitations of ChatGPT for multimedia forensic tasks.

深度伪造大模型音视频检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。