轻量级无监督模型统一解决语音增强与分离问题
Lightweight Wasserstein Audio-Visual Model for Unified Speech Enhancement and Separation
- 利用唇动和面部特征引导语音提取,无需配对数据
- 采用Wasserstein距离正则化,稳定潜在空间表现
- 在噪声和多人说话场景下均表现优异,适合实际应用
语音增强(SE)和语音分离(SS)传统上被视为独立任务。然而,真实场景中常同时存在背景噪声和多说话人重叠,亟需统一解决方案。现有方法虽尝试在多阶段架构中融合两者,但普遍依赖复杂、参数量大的模型及有监督训练,限制了可扩展性与泛化能力。本文提出UniVoiceLite,一种轻量级、无监督的音视频统一框架,将SE与SS整合于单一模型中。该模型利用唇部运动与面部身份线索指导语音提取,并通过Wasserstein距离正则化稳定潜在空间,无需成对的噪声-清晰语音数据。实验表明,UniVoiceLite在噪声环境与多说话人场景下均表现出色,兼具高效性与强泛化能力。代码已开源:https://github.com/jisoo-o/UniVoiceLite。
原文摘要 · Abstract (English)
Speech Enhancement (SE) and Speech Separation (SS) have traditionally been treated as distinct tasks in speech processing. However, real-world audio often involves both background noise and overlapping speakers, motivating the need for a unified solution. While recent approaches have attempted to integrate SE and SS within multi-stage architectures, these approaches typically involve complex, parameter-heavy models and rely on supervised training, limiting scalability and generalization. In this work, we propose UniVoiceLite, a lightweight and unsupervised audio-visual framework that unifies SE and SS within a single model. UniVoiceLite leverages lip motion and facial identity cues to guide speech extraction and employs Wasserstein distance regularization to stabilize the latent space without requiring paired noisy-clean data. Experimental results demonstrate that UniVoiceLite achieves strong performance in both noisy and multi-speaker scenarios, combining efficiency with robust generalization. The source code is available at https://github.com/jisoo-o/UniVoiceLite.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。