arXiv:2605.28064eess.AScs.AI2026-05

人类对语音伪造的识别依赖感知与情境,而非单纯技术判断。

I Hear, Therefore I Trust: A Socio-Technical Investigation of Humans as Synthetic Speech Detectors

论文配图:I Hear, Therefore I Trust: A Socio-Technical Investigation of Humans as Synthetic Speech Detectors
图 1 · 摘自论文原文
  • 通过47人实验,分析人在不同信任线索下定位语音伪造片段的能力。
  • 完全合成语音的检测准确率低于随机水平,但质量评分能反映真实类型。
  • 适合关注人机交互、虚假信息防御的研究者阅读。

自动深度伪造检测受到广泛关注,但人类实际遭遇合成语音的社会技术环境仍不清晰。本研究将语音深度伪造检测视为一种感知与情境过程,设计了一项定位任务:47名参与者在三种受控信任线索(指令框架、情感预示、来源标注)下,标记真实、完全合成及部分合成语句中疑似伪造的片段。参与者还对机械感、表现力、可理解性、清晰度、镇定度和评估信心进行评分。语句类型是影响检测准确性和感知质量的主要因素;信任线索未产生主效应,但驱动了检测行为。完全合成语音的检测准确率低于随机水平,而质量评分与语句类型高度相关,表明即使公开检测失败,仍存在隐性区分。

原文摘要 · Abstract (English)

Automatic deepfake detection has received considerable research attention, yet the socio-technical environment in which humans actually encounter synthetic speech remains poorly understood. We investigate voice deepfake detection as a perceptual and contextual process, presenting a localization task in which 47 participants marked suspected synthetic segments across authentic, fully synthetic, and partially synthetic utterances under three manipulated trust cues: instructional framing, affective priming, and provenance labeling. Participants provided quality ratings on mechanicalness, expressiveness, intelligibility, clarity, calmness, and confidence of evaluation. Utterance class was the primary determinant of detection accuracy and perceptual quality; trust cues produced no main effects but motivated detection behavior. Fully synthetic speech was detected at below-chance levels. Quality ratings tracked utterance type, indicating implicit discrimination where overt detection failed.

语音伪造人因检测信任机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。