arXiv:2608.07980eess.AScs.CL2026-08

语音不是唯一生物特征,别再用‘声纹’当身份凭证了。

The Voiceprint Fallacy: Why Voices Are Not Unique Biometric Imprints

论文配图:The Voiceprint Fallacy: Why Voices Are Not Unique Biometric Imprints
图 1 · 摘自论文原文
  • 驳斥‘声纹唯一’神话,指出语音具高度可变性
  • 实证显示语音受语境、情绪等影响极大,无法稳定识别
  • 适合关注语音安全、身份验证的开发者与政策制定者

近年来,'声纹'一词重新受到技术与政策领域关注,常被视作类似指纹的稳定唯一生物特征。然而,自其提出以来,法医语音专家屡次批评这一概念。尽管语音确实包含说话人信息,但将语音视为固定不变的身份标记,掩盖了语言的高度动态与情境依赖性。本文回顾声纹谬误的历史,分析人类语音变异、法医语音比对、人工与人类语音识别研究,以及深度伪造语音对身份认定的挑战。指出'声纹'隐喻在科学上具有误导性,它将概率性语音线索误构为稳定身份实体。我们主张,当前证据不支持存在稳定且个体唯一的声纹。对于语音识别与生物识别应用,应基于训练与评估条件理解学习到的说话人表征,并明确评估其对说话人内部变化、领域差异及合成伪造的鲁棒性。

原文摘要 · Abstract (English)

In recent years, the term voiceprint has regained attention, particularly in technological applications and policy-making contexts, often carrying the assumption that a person's voice constitutes a stable and unique biometric trace analogous to a fingerprint. Yet this conception has been repeatedly criticized and rejected by forensic voice experts throughout the decades since its introduction. Although voices undoubtedly contain speaker-related information, this simplified conception obscures the highly dynamic and context-dependent nature of speech. This article revisits the voiceprint fallacy and reconsiders what can count as evidence of speaker identity by reviewing the historical development of voiceprint identification, evidence on human voice variability, developments in forensic voice comparison, research on human and automatic speaker recognition, and the recent challenge posed by deepfake speech to speaker identity. We point out that the voiceprint metaphor and its underlying implications are scientifically misleading because they transform a probabilistic source of speaker information into an imagined stable object of identity. We argue that speaker identity assessment does not require, and current evidence does not support, the existence of a stable and individually unique voiceprint. For speaker recognition and voice biometrics, this distinction motivates interpreting learned speaker representations with respect to the conditions under which they are trained and evaluated, and explicitly assessing their robustness to relevant sources of within-speaker variability, domain mismatch, and synthetic manipulation.

语音识别生物识别深度伪造身份验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。