arXiv:2511.21580cs.SD2025-11

用分离音色的编码器+Transformer预测高频,提升音频宽带扩展质量

Harmonic-Percussive Disentangled Neural Audio Codec for Bandwidth Extension

  • 基于谐波与打击乐分离的神经编码器,将音频转为离散符号
  • 用Transformer预测缺失高频部分,主观听感和客观指标均优于现有方法
  • 适合做音频修复、老歌增强的工程师或研究者参考

宽带扩展是将低通音频信号中的高频成分重建出来,属于音频处理中的长期难题。本文将该任务视为音频令牌预测问题,采用基于Transformer的语言模型,在由谐波-打击乐分解引导的解耦神经音频编码器生成的离散表示上进行训练。该编码器显式考虑下游预测任务,实现编码结构与变压器建模的高效耦合。实验表明,该联合设计在客观指标和主观评估中均获得高质量重建效果。结果强调了将编码解耦与表征学习与生成建模阶段对齐的重要性,并展示了全局、表征感知设计在推进宽带扩展中的潜力。

原文摘要 · Abstract (English)

Bandwidth extension, the task of reconstructing the high-frequency components of an audio signal from its low-pass counterpart, is a long-standing problem in audio processing. While traditional approaches have evolved alongside the broader trends in signal processing, recent advances in neural architectures have significantly improved performance across a wide range of audio tasks, In this work, we extend these advances by framing bandwidth extension as an audio token prediction problem. Specifically, we train a transformer-based language model on the discrete representations produced by a disentangled neural audio codec, where the disentanglement is guided by a Harmonic-Percussive decomposition of the input signals, highlighting spectral structures particularly relevant for bandwidth extension. Our approach introduces a novel codec design that explicitly accounts for the downstream token prediction task, enabling a more effective coupling between codec structure and transformer modeling. This joint design yields high-quality reconstructions of the original signal, as measured by both objective metrics and subjective evaluations. These results highlight the importance of aligning codec disentanglement and representation learning with the generative modeling stage, and demonstrate the potential of global, representation-aware design for advancing bandwidth extension.

音频修复Transformer编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。