arXiv:2510.13221eess.AS2025-10被引 1

分离语音内容与环境声学特征,实现语音的声学空间传送。

Acoustic Teleportation via Disentangled Neural Audio Codec Representations

  • 通过解耦语音内容与房间声学特征,实现跨场景语音传送。
  • 非侵入式评分达3.03,显著优于之前方法的2.44。
  • 嵌入向量能准确反映房间混响时间,适合语音增强与虚拟会议。

本文提出一种基于神经音频编解码器表示的声学传送方法,通过将语音内容与声学环境特征解耦,实现语音录音间房间特性的迁移,同时保持内容和说话人身份不变。基于EnCodec架构,采用五项训练任务:纯净重建、混响重建、去混响及两种变体的声学传送。实验显示,对声学嵌入进行时序下采样会显著降低性能,即使2倍下采样也导致质量统计上显著下降。学习到的声学嵌入与RT60存在强相关性。t-SNE聚类分析表明,声学嵌入按房间聚类,语音嵌入按说话人聚类,验证了有效解耦。

原文摘要 · Abstract (English)

This paper presents an approach for acoustic teleportation by disentangling speech content from acoustic environment characteristics in neural audio codec representations. Acoustic teleportation transfers room characteristics between speech recordings while preserving content and speaker identity. We build upon previous work using the EnCodec architecture, achieving substantial objective quality improvements with non-intrusive ScoreQ scores of 3.03, compared to 2.44 for prior methods. Our training strategy incorporates five tasks: clean reconstruction, reverberated reconstruction, dereverberation, and two variants of acoustic teleportation. We demonstrate that temporal downsampling of the acoustic embedding significantly degrades performance, with even 2x downsampling resulting in a statistically significant reduction in quality. The learned acoustic embeddings exhibit strong correlations with RT60. Effective disentanglement is demonstrated using t-SNE clustering analysis, where acoustic embeddings cluster by room while speech embeddings cluster by speaker.

声学传送语音处理嵌入解耦环境建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。