人类难辨音视频深度伪造,但可集体筛查真伪。
Beyond Seeing Is Believing: On Crowdsourced Detection of Audiovisual Deepfakes

- 用众包测试人类识别真假音视频的能力。
- 96个视频中多数伪造被漏检,且对篡改类型识别不准。
- 适合用于大规模真实性初筛,但无法可靠判断篡改方式。
深度伪造日益逼真且易制作,引发人们对误信息环境中人类判断可靠性担忧。本文通过在Prolific平台开展两项匹配的众包实验,基于AV-Deepfake1M和可信媒体挑战(TMC)数据集,每数据集抽取48个视频(共96个),每个视频收集10份判断(总计960份)。结果表明:人群极少将真实视频误判为伪造,但普遍漏检伪造内容,跨视频一致性有限。聚合多份判断可稳定真实性信号,但无法恢复多数人一致遗漏的伪造。即使发现伪造,篡改类型识别误差较大,尤其是音视频联合篡改情况最难识别。总体表明,众包可提供可扩展的音视频真实性初步筛查信号,但可靠的模态归因仍是开放挑战。
原文摘要 · Abstract (English)
Deepfakes are increasingly realistic and easy to produce, raising concerns about the reliability of human judgments in misinformation settings. We study audiovisual deepfake detection by measuring how consistently crowd workers distinguish authentic from manipulated videos and, when they flag a video as manipulated, how accurately they identify the manipulation type (audio-only, video-only, or audio-video) and how consistently they report manipulation timestamps. We run two matched crowdsourcing studies on Prolific using AV-Deepfake1M and the Trusted Media Challenge (TMC) dataset. We sample 48 videos per dataset (96 total) and collect 960 judgments (10 per video). Results show that crowd workers rarely misclassify authentic videos as manipulated, but they miss many manipulations, and agreement remains limited across videos. Aggregating multiple judgments per video stabilizes the authenticity signal, but it cannot recover manipulations that most workers consistently miss. Manipulation type identification is substantially noisier than authenticity detection even when workers detect a manipulation, with joint audio-video cases being particularly hard to recognize. Overall, these findings suggest that crowdsourcing can provide a scalable screening signal for audiovisual authenticity, while reliable modality attribution remains an open challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。