arXiv:2512.16304cs.SD2025-12

用思维链引导语音超分辨率,让低采样音频恢复更准确

CogSR: Semantic-Aware Speech Super-Resolution via Chain-of-Thought Guided Flow Matching

  • 通过思维链推理引入语义锚点,指导音频重建
  • 在严重退化情况下仍能还原真实高频细节与语言内容
  • 适合历史档案和监控音频等高价值低质量音频修复

将语音超分辨率(SR)应用于严重低采样率录音是数字存档和调查音频恢复中的关键挑战。此类输入缺乏必要声学线索,导致现有生成模型常因上下文不足而产生错误语音内容,仅基于概率猜测词汇。为此,我们提出 CogSR 框架,专为高精度离线修复设计。方法从简单信号映射转向认知重建:结合大音频语言模型,利用思维链(Chain-of-Thought)推理作为语义锚点,同时通过显式声学先验保持说话人身份一致。该机制引导修正流(Rectified Flow)骨干网络合成不仅逼真且语言准确的高频细节。实验表明,CogSR 能有效消除严重退化场景下的歧义,成为恢复高价值遗留与监控音频的稳健解决方案。

原文摘要 · Abstract (English)

Applying speech super-resolution (SR) to recordings with severely low sampling rates is a critical challenge in digital archiving and investigative audio recovery. In these scenarios, the input lacks essential acoustic cues. Consequently, existing generative models often fail; without sufficient context, they hallucinate phonetic content, guessing words based on probability rather than meaning. To address this, we propose CogSR, a framework designed specifically for high-precision, offline restoration. Our approach shifts the focus from simple signal mapping to cognitive reconstruction. By integrating a Large Audio-Language Model, we employ Chain-of-Thought reasoning to act as a semantic anchor, while explicit acoustic priors ensure the speaker's identity remains consistent. This guides a Rectified Flow backbone to synthesize high-frequency details that are not only realistic but linguistically accurate. Evaluations show that CogSR effectively eliminates ambiguity in severe degradation regimes, making it a robust solution for restoring high-value legacy and surveillance audio.

语音修复思维链音频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。