arXiv:2508.02391cs.SDcs.AI2025-08AAAI被引 5

通过推理时扩展提升音频超分质量,突破传统采样限制。

Inference-time Scaling for Diffusion-based Audio Super-resolution

  • 引入推理时扩展机制,多路径搜索优化采样过程。
  • 语音超分在4kHz到24kHz下提升46.98%频谱距离、15.20%词错误率。
  • 适用于影视后期、音乐母带等高算力场景的高质量音频生成。

扩散模型在生成任务中表现卓越,包括音频超分辨率(SR)。在电影后期制作和专辑母带处理等场景中,有充足计算资源可用于追求更高音频质量。然而,现有扩散方法通常通过增加采样步数来提升质量,其性能仍受限于采样过程的随机性,导致输出方差大、质量上限低。本文提出一种新的推理时扩展范式,通过在采样过程中探索多个解路径,设计特定任务验证器,并引入随机搜索与零阶搜索算法,结合验证器-算法组合主动引导高维解空间的探索,实现更鲁棒、更高品质的输出。在多种音频领域(语音、音乐、音效)及频率范围的广泛验证中,均取得稳定提升:语音超分从4kHz到24kHz,美学评分提升9.70%,说话人相似度提升5.88%,词错误率降低15.20%,频谱距离减少46.98%。音频样本可访问:https://racerk.github.io/tt-scale-audiosr/。

原文摘要 · Abstract (English)

Diffusion models have demonstrated remarkable success in generative tasks, including audio super-resolution (SR). In many applications like movie post-production and album mastering, substantial computational budgets are available for achieving superior audio quality. However, while existing diffusion approaches typically increase sampling steps to improve quality, the performance remains fundamentally limited by the stochastic nature of the sampling process, leading to high-variance and quality-limited outputs. Here, rather than simply increasing the number of sampling steps, we propose a different paradigm through inference-time scaling for SR, which explores multiple solution trajectories during the sampling process. Different task-specific verifiers are developed, and two search algorithms, including the random search and zero-order search for SR, are introduced. By actively guiding the exploration of the high-dimensional solution space through verifier-algorithm combinations, we enable more robust and higher-quality outputs. Through extensive validation across diverse audio domains (speech, music, sound effects) and frequency ranges, we demonstrate consistent performance gains, achieving improvements of up to 9.70% in aesthetics, 5.88% in speaker similarity, 15.20% in word error rate, and 46.98% in spectral distance for speech SR from 4kHz to 24kHz, showcasing the effectiveness of our approach. Audio samples are available at: https://racerk.github.io/tt-scale-audiosr/.

音频超分扩散模型推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。