arXiv:2606.31365eess.AScs.SD2026-06中稿 · Interspeech 2026

用探针评估音频编码器解耦效果,发现声学信息会泄露到语音嵌入中。

Beyond Cross-Reconstruction: Probing-Based Disentanglement Evaluation for Acoustic Teleportation Codecs

论文配图:Beyond Cross-Reconstruction: Probing-Based Disentanglement Evaluation for Acoustic Teleportation Codecs
图 1 · 摘自论文原文
  • 通过回归房间参数和分类说话人来检测嵌入空间解耦程度
  • 说话人身份基本保留在独立分区,但声学特征存在泄漏
  • 无需显式监督,声学嵌入已能精准估计房间特性

一些神经音频编码器将语音分解为内容、说话人身份和声学等潜在子空间,支持声学传送和语音转换。现有评估依赖交叉重建质量,无法可靠检测各分区间的泄漏。本文扩展基于探针的框架,通过回归房间声学参数(混响时间、清晰度、直达比)和分类说话人身份,以有意与无意分区间的差距作为解耦度量。应用于声学传送编码器发现,说话人身份基本保留在其分区,而声学特征因训练目标导致向语音嵌入泄漏。声学嵌入在0.02秒内即达到监督基线的房间参数估计精度,表明物理意义结构在无显式监督下自然形成。

原文摘要 · Abstract (English)

Some neural audio codecs disentangle speech into latent subspaces encoding content, speaker identity, and acoustics, enabling acoustic teleportation and voice conversion. Existing evaluations rely on cross-reconstruction quality, which cannot reliably detect leakage across partitions. We extend a probing based framework to assess disentanglement by regressing room-acoustic parameters (reverberation time, clarity, and direct-to-reverberant ratio) and classifying speaker identity, using the gap between intended and unintended partitions as the disentanglement measure. Applied to an acoustic teleportation codec, we find speaker identity is largely confined to its partition, while acoustics leak into the speech embeddings due to the training objective. Acoustic embeddings blindly estimate room parameters within 0.02 s of supervised baselines, indicating physically meaningful structure emerges without explicit supervision.

音频编码解耦表征探针评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。