arXiv:2510.05718eess.AS2025-10被引 1

发现说话人嵌入中机器与人类感知不一致,实现完全保留人耳感知的语音匿名化。

Investigation of perception inconsistency in speaker embedding for asynchronous voice anonymization

  • 通过修改说话人嵌入,揭示其内部存在影响机器但不影响人类感知的子空间。
  • 在FACodec和Diff-HierVC模型上验证,实现100%人类感知保留率。
  • 适合关注语音隐私保护与感知一致性研究的研究者。

在基于嵌入向量表示说话人属性的语音生成框架中,可通过修改原始语音提取的说话人嵌入实现异步语音匿名化。然而,说话人嵌入中机器与人类对说话人属性的感知不一致尚未被探索,限制了异步语音匿名化的效果。本研究通过在语音生成过程中修改说话人嵌入,探究这一不一致现象。在FACodec和Diff-HierVC语音生成模型上的实验发现,存在一个子空间:移除该子空间可改变机器对说话人属性的判断,却保持人类对其感知不变。基于此发现,提出一种异步语音匿名化方法,在保持100%人类感知保留率的同时有效遮蔽机器感知。音频样本见 https://voiceprivacy.github.io/speaker-embedding-eigen-decomposition/。

原文摘要 · Abstract (English)

Given the speech generation framework that represents the speaker attribute with an embedding vector, asynchronous voice anonymization can be achieved by modifying the speaker embedding derived from the original speech. However, the inconsistency between machine and human perceptions of the speaker attribute within the speaker embedding remains unexplored, limiting its performance in asynchronous voice anonymization. To this end, this study investigates this inconsistency via modifications to speaker embedding in the speech generation process. Experiments conducted on the FACodec and Diff-HierVC speech generation models discover a subspace whose removal alters machine perception while preserving its human perception of the speaker attribute in the generated speech. With these findings, an asynchronous voice anonymization is developed, achieving 100% human perception preservation rate while obscuring the machine perception. Audio samples can be found in https://voiceprivacy.github.io/speaker-embedding-eigen-decomposition/.

语音匿名化说话人嵌入感知一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。