用离散高分辨率音频标记实现高质量语音增强
High-Fidelity Speech Enhancement via Discrete Audio Tokens
- 基于离散音频标记构建简化语言模型框架
- 在客观指标和主观评测中均超越现有方法
- 适合追求高保真语音增强的研究与应用
近期基于自回归Transformer的语音增强方法借助先进的语义理解和上下文建模取得良好效果,但通常依赖复杂的多阶段流程和低采样率编码器,限制了其通用性。本文提出DAC-SE1,一种基于离散高分辨率音频表示的简化语言模型框架,在保留精细声学细节的同时保持语义连贯性。实验表明,DAC-SE1在客观感知指标和MUSHRA主观评估中均优于现有自回归语音增强方法。代码与模型权重已开源,以支持可扩展、统一且高质量的语音增强研究。
原文摘要 · Abstract (English)
Recent autoregressive transformer-based speech enhancement (SE) methods have shown promising results by leveraging advanced semantic understanding and contextual modeling of speech. However, these approaches often rely on complex multi-stage pipelines and low sampling rate codecs, limiting them to narrow and task-specific speech enhancement. In this work, we introduce DAC-SE1, a simplified language model-based SE framework leveraging discrete high-resolution audio representations; DAC-SE1 preserves fine-grained acoustic details while maintaining semantic coherence. Our experiments show that DAC-SE1 surpasses state-of-the-art autoregressive SE methods on both objective perceptual metrics and in a MUSHRA human evaluation. We release our codebase and model checkpoints to support further research in scalable, unified, and high-quality speech enhancement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。