arXiv:2608.00572cs.SD2026-08

一个模型搞定各种带宽的语音扩频,无需重新训练。

AnyBand: Unified Multi-Bandwidth Speech Extension via Frequency-Aware In-Context Spectral Infilling

论文配图:AnyBand: Unified Multi-Bandwidth Speech Extension via Frequency-Aware In-Context Spectral Infilling
图 1 · 摘自论文原文
  • 用低频谱作提示,实时补全高频内容
  • 在多个数据集上重建精度优于现有方法
  • 适合实际中带宽不固定的语音处理场景

语音带宽扩展(BWE)旨在从带宽受限的语音中恢复缺失的高频信息。现有方法通常将BWE视为固定或预定义的带宽转换问题,可能需要针对不同截止频率设计专用模型或重新训练,限制了其在真实场景中的适用性。本文提出AnyBand,一种统一的BWE框架,将带宽扩展重构为上下文相关的谱图补全任务。受提示学习启发,AnyBand以观测到的低频谱作为频域提示,传递内容、说话人、语调和频谱包络等线索,实现单一模型对连续输入带宽范围的截断条件生成。模型通过缺失频段条件流匹配与连续采样的截止频率课程学习进行训练。为更好利用频谱提示,引入频率感知扩散变换器,建模跨频段交互与长时依赖,并采用物理驱动的多视角对抗精修阶段提升频谱真实感、包络连贯性与谐波一致性。在多个数据集和带宽设置下实验表明,AnyBand在谱重建上持续优于基线方法,且在标准与非规则输入截止频率下均达到竞争性听觉质量。音频样本可获取。

原文摘要 · Abstract (English)

Bandwidth extension (BWE) aims to recover missing high-frequency content from band-limited speech. Existing methods often formulate BWE as a fixed or predefined bandwidth conversion problem, potentially requiring cutoff-specific models or retraining when the input bandwidth changes. This assumption limits their applicability to practical scenarios where speech may arrive with diverse cutoff frequencies. We propose AnyBand, a unified BWE framework that recasts bandwidth extension as in-context spectral infilling. Motivated by prompt-based zero-shot speech generation, AnyBand conditions high-frequency generation on the observed low-frequency spectrum, using the available band as a frequency-domain prompt that conveys content, speaker, prosodic, and spectral-envelope cues. This formulation enables a single model to perform cutoff-conditioned generation over a continuous range of input bandwidths. AnyBand is trained with missing-band conditional flow matching and an Easy-to-Balanced cutoff curriculum over continuously sampled cutoff frequencies. To better exploit the spectral prompt, we introduce a frequency-aware Diffusion Transformer that models cross-frequency interactions and long-range temporal dependencies, followed by a physically motivated multi-view adversarial refinement stage to enhance spectral realism, envelope coherence, and harmonic consistency. Experiments on multiple datasets and bandwidth settings show that AnyBand consistently improves spectral reconstruction over existing baselines while achieving competitive perceptual quality across both standard and irregular input cutoffs. Audio samples are available.

语音扩展频谱补全扩散模型统一模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。