arXiv:2607.16870cs.SDcs.AI2026-07

语音令牌可能泄露声音特征,攻击者可借此还原说话人身份。

Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models

论文配图:Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models
图 1 · 摘自论文原文
  • 用可训练模型构建语音令牌表征,捕捉说话人敏感信息。
  • 仅需3秒输出,即可在指定编码器空间实现0.7以上相似度还原。
  • 适合关注语音隐私安全的研究者与产品开发者。

端到端语音语言模型越来越多地使用语音令牌表示用户语音,而非依赖传统的ASR--LLM--TTS流水线。尽管这些令牌支持更自然、低延迟的语音交互,但也可能保留敏感的说话人特征。我们研究了暴露的语音令牌是否会泄露声音特征,并将其风险定义为说话人逆向攻击。提出Audio BERT(AuB),一种从离散代码本构建令牌嵌入并聚合为说话人敏感表征的可训练模型;并设计SpInv,一种基于AuB的两阶段逆向方法,用于在攻击者指定的说话人编码器空间中恢复嵌入。在VoxCeleb数据集上,采用说话人独立协议评估Moshi、Higgs3、Kimi-Audio和Qwen3-Omni模型。大量实验表明,仅需三秒前端输出,SpInv在攻击者指定的说话人编码器空间中达到超过0.70的余弦相似度。

原文摘要 · Abstract (English)

End-to-end speech language models increasingly represent user speech with speech tokens rather than relying exclusively on cascaded ASR--LLM--TTS pipelines. Although these tokens support expressive and low-latency spoken interaction, they may also preserve sensitive speaker characteristics. We investigate whether exposed speech tokens leak voiceprints and formulate this risk as a speaker inversion attack. We introduce Audio BERT (AuB), a trainable model that constructs token embeddings from discrete codebooks and aggregates them into speaker-sensitive representations, and propose SpInv, a two-stage inversion method built on AuB to recover embeddings in the space of an attacker-specified speaker encoder. We evaluate Moshi, Higgs3, Kimi-Audio, and Qwen3-Omni using speaker-disjoint protocols on the VoxCeleb dataset. Extensive experiments show that, with only three seconds of frontend output, SpInv achieves cosine similarities above 0.70 in the attacker-specified speaker-encoder space.

语音安全隐私泄露说话人识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。