arXiv:2509.18579eess.AScs.CL2025-09被引 2

用文本模型教音频模型推理,提升复杂任务表现。

Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation

  • 双维度知识蒸馏:跨模态与分层对齐
  • 音频推理性能显著提升,保留声学能力
  • 适合需要强推理的语音任务研究者

尽管大型音频语言模型在自动语音识别和情感识别等任务上表现优异,但因音频与文本之间的模态差异以及缺乏结构化中间监督,仍难以处理复杂推理。为此,我们提出一种统一的知识蒸馏框架,将高容量文本教师模型的推理能力迁移至学生音频模型,同时保持其声学建模能力。方法引入两个关键维度:源级蒸馏,利用文本和声学教师提供互补的模态特异性监督;层级蒸馏,将教师信号与对应的学生层对齐,提升迁移效率。该双重策略实现对蒸馏过程的细粒度控制,有效弥合符号推理与语音表征之间的差距。实验表明,音频推理性能显著提升,验证了该框架作为音频建模中推理迁移方案的有效性。

原文摘要 · Abstract (English)

While large audio language models excel at tasks like ASR and emotion recognition, they still struggle with complex reasoning due to the modality gap between audio and text as well as the lack of structured intermediate supervision. To address this, we propose a unified knowledge distillation framework to transfer reasoning capabilities from a high-capacity textual teacher model to a student audio models while preserving its acoustic competence. Our method introduces two key dimensions: source-wise distillation, which leverages both textual and acoustic teachers to provide complementary modality-specific supervision; and layer-wise distillation, which aligns teacher signals with appropriate student layers to improve transfer efficiency. This dual-dimensional strategy enables fine-grained control over the distillation process, effectively bridging the gap between symbolic reasoning and speech representations. Experimental results show significant improvements in audio reasoning performance, demonstrating the effectiveness of our framework as a reasoning transfer solution for audio modeling.

音频模型知识蒸馏推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。