从杂乱对话音频中还原清晰的双人对话轨道,提升可懂度与分离效果。
DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio

- 用自编码器压缩语音特征,再通过扩散模型恢复说话人独立信号。
- 在真实场景数据上显著提升可懂度与分离质量,推理速度更快。
- 适合语音分离、对话分析和语音增强领域的研究者使用。
全双工对话音频(每位说话人独立录音)对口语对话研究至关重要,但难以大规模获取。大多数真实场景下的双人对话仅以退化的单声道混合音频形式存在,不适合需要纯净说话人信号的系统。我们提出DialogueSidon,一种联合恢复与分离退化单声道双人对话音频的模型。DialogueSidon结合了基于语音自监督学习(SSL)特征的变分自编码器(VAE),将SSL特征压缩至紧凑潜在空间,并利用基于扩散的潜在预测器从退化混合信号中恢复说话人级潜在表示。在英文、多语言及真实场景对话数据集上的实验表明,DialogueSidon在可懂度和分离质量上均显著优于基线模型,同时实现更快的推理速度。
原文摘要 · Abstract (English)
Full-duplex dialogue audio, in which each speaker is recorded on a separate track, is an important resource for spoken dialogue research, but is difficult to collect at scale. Most in-the-wild two-speaker dialogue is available only as degraded monaural mixtures, making it unsuitable for systems requiring clean speaker-wise signals. We propose DialogueSidon, a model for joint restoration and separation of degraded monaural two-speaker dialogue audio. DialogueSidon combines a variational autoencoder (VAE) operates on the speech self-supervised learning (SSL) model feature, which compresses SSL model features into a compact latent space, with a diffusion-based latent predictor that recovers speaker-wise latent representations from the degraded mixture. Experiments on English, multilingual, and in-the-wild dialogue datasets show that DialogueSidon substantially improves intelligibility and separation quality over a baseline, while also achieving much faster inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。