arXiv:2512.09221eess.ASeess.SP2025-12被引 3

研究人类如何辨别语音深度伪造,发现语调和说话风格是关键线索。

Human perception of audio deepfakes: the role of language and speaking style

  • 通过54名听者实验,测试语言、语体和熟悉度对识别真假语音的影响。
  • 平均识别准确率59.11%,对真实语音判断更准,依赖语调节奏等非语音特征。
  • 揭示跨语言差异,适合语音安全与人机交互研究者阅读。

语音深度伪造已达到难以区分真伪的程度,带来身份盗用和虚假信息传播风险。尽管如此,关于人类识别能力的研究仍有限,多数聚焦英语,且缺乏对判断依据的深入分析。本研究通过感知实验,让54名听者(28名西语母语者,26名日语母语者)对80个语音样本(50%为人工合成)进行自然/合成分类并说明理由。样本按语言(西语/日语)、语体(有声书/访谈)和熟悉度(熟悉/不熟悉)分组。结果表明,平均准确率为59.11%,对真实语音的识别表现更优。判断主要依赖语言与非语言线索的结合,尤其重视语调、节奏、流畅性、停顿、语速、呼吸和笑声等超音段特征。定性分析显示,日语与西语听者在“语音人性感”认知上既有共性也有跨语言差异。研究结果支持先前关于韵律和自发语流特征重要性的观点,揭示了人类辨别语音真伪的复杂机制。

原文摘要 · Abstract (English)

Audio deepfakes have reached a level of realism that makes it increasingly difficult to distinguish between human and artificial voices, which poses risks such as identity theft or spread of disinformation. Despite these concerns, research on humans' ability to identify deepfakes is limited, with most studies focusing on English and very few exploring the reasons behind listeners' perceptual decisions. This study addresses this gap through a perceptual experiment in which 54 listeners (28 native Spanish speakers and 26 native Japanese speakers) classified voices as natural or synthetic, and justified their choices. The experiment included 80 stimuli (50% artificial), organized according to three variables: language (Spanish/Japanese), speech style (audiobooks/interviews), and familiarity with the voice (familiar/unfamiliar). The goal was to examine how these variables influence detection and to analyze qualitatively the reasoning behind listeners' perceptual decisions. Results indicate an average accuracy of 59.11%, with higher performance on authentic samples. Judgments of vocal naturalness rely on a combination of linguistic and non-linguistic cues. Comparing Japanese and Spanish listeners, our qualitative analysis further reveals both shared cues and notable cross-linguistic differences in how listeners conceptualize the "humanness" of speech. Overall, participants relied primarily on suprasegmental and higher-level or extralinguistic characteristics - such as intonation, rhythm, fluency, pauses, speed, breathing, and laughter - over segmental features. These findings underscore the complexity of human perceptual strategies in distinguishing natural from artificial speech and align partly with prior research emphasizing the importance of prosody and phenomena typical of spontaneous speech, such as disfluencies.

语音伪造人类感知跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。