arXiv:2602.01284cs.MMcs.CV2026-02被引 2

研究人类识别深度伪造视频的多模态策略,助力提升媒介素养。

Seeing, Hearing, and Knowing Together: Multimodal Strategies in Deepfake Videos Detection

  • 通过视觉、听觉和直觉线索联合判断真假视频
  • 真实视频识别准确率高于伪造视频,且信心更匹配实际表现
  • 强调多模态线索组合是有效识别的关键,适合设计防伪教育工具

随着深度伪造视频日益逼真,理解人类识别策略对设计有效的媒介素养干预措施至关重要。本研究招募了195名21至40岁参与者,评估他们对真实与伪造视频的判断,记录其置信度及依赖的视觉、音频和知识类线索。结果显示,参与者在真实视频上的识别准确率更高,且对真实内容的期望校准误差更低。通过关联规则挖掘,发现成功识别常伴随视觉外观、语音特征和直觉线索的共同使用,凸显多模态策略的重要性。研究揭示了哪些线索有助于或阻碍识别,为设计引导有效线索使用的媒介素养工具提供了方向,有助于提升公众对虚假数字媒体的抵御能力。

原文摘要 · Abstract (English)

As deepfake videos become increasingly difficult for people to recognise, understanding the strategies humans use is key to designing effective media literacy interventions. We conducted a study with 195 participants between the ages of 21 and 40, who judged real and deepfake videos, rated their confidence, and reported the cues they relied on across visual, audio, and knowledge strategies. Participants were more accurate with real videos than with deepfakes and showed lower expected calibration error for real content. Through association rule mining, we identified cue combinations that shaped performance. Visual appearance, vocal, and intuition often co-occurred for successful identifications, which highlights the importance of multimodal approaches in human detection. Our findings show which cues help or hinder detection and suggest directions for designing media literacy tools that guide effective cue use. Building on these insights can help people improve their identification skills and become more resilient to deceptive digital media.

深度伪造多模态媒介素养

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。