arXiv:2607.19636eess.AS2026-07

多轮语音融合可突破语音匿名化,暴露说话人身份。

Multimodal Speaker Verification as a Threat to Speaker Anonymization

  • 融合多轮音频、韵律和语言信息进行身份验证
  • 仅5段匿名语音即可使错误率降低15%以上
  • 帧级聚合效果最佳,适合隐私保护研究者

大多数自动说话人验证(ASV)系统基于单次语音片段,但现实交互通常包含多轮语音。随着语音积累,声学、语调和语言线索提供的说话人信息逐渐丰富,可能威胁仅针对语音特征设计的匿名化方法。本文在多轮、多模态设置下研究ASV,考察跨匿名语音的信息聚合是否影响隐私。首先分析仅音频的多轮聚合,发现随着语音数量增加,性能持续提升。随后引入语调和语言信息,显示多模态系统优于单模态方法。最后比较聚合策略,发现帧级聚合达到最低等错误率(EER)。即使仅有5段匿名语音,音频与文本结合仍比纯音频聚合降低超15%的EER,表明即便经过匿名化,仍存在大量可识别的说话人信息。

原文摘要 · Abstract (English)

Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions typically consisting of multiple utterances. As speech accumulates, increasingly rich speaker information becomes available through acoustic, prosodic, and linguistic cues, potentially challenging speaker anonymization methods that primarily target vocal characteristics. We investigate ASV in a multi-utterance, multimodal setting and examine whether aggregating information across anonymized speech impacts privacy. We first study audio-only aggregation across multiple anonymized utterances and observe consistent performance improvements as more speech becomes available. We then incorporate prosodic and linguistic information, showing that multimodal systems outperform unimodal approaches. Finally, we compare aggregation strategies and find that frame-level aggregation yields the lowest EERs. Even with only five anonymized utterances, combining audio and text reduces EER by over 15% relative to audio-only aggregation, demonstrating that substantial speaker-discriminative information remains accessible despite anonymization.

说话人验证隐私保护多模态匿名化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。