arXiv:2509.14930cs.CLcs.AI2025-09被引 7

用跨模态知识蒸馏,让语音大模型不丢文本知识还能更好理解口语。

Cross-Modal Knowledge Distillation for Speech Large Language Models

  • 通过文本-文本与语音-文本双通道蒸馏,迁移文本模型知识到语音模型。
  • 在对话和语音理解任务中,保留了90%以上的文本知识,推理能力提升显著。
  • 适合研究语音大模型、跨模态对齐或想避免知识遗忘的开发者。

本文首次系统评估了语音大语言模型中的灾难性遗忘与模态不平等问题,发现引入语音能力会损害文本知识与推理能力,即使输入仍为文本;且在语音查询时性能进一步下降。为此,我们提出一种跨模态知识蒸馏框架,利用文本-文本与语音-文本双通道,将文本教师模型的知识迁移到语音大模型。在对话与音频理解任务上的大量实验表明,该方法能有效保留文本知识,改善跨模态对齐,并增强基于语音的交互推理能力。

原文摘要 · Abstract (English)

In this work, we present the first systematic evaluation of catastrophic forgetting and modality inequivalence in speech large language models, showing that introducing speech capabilities can degrade knowledge and reasoning even when inputs remain textual, and performance further decreases with spoken queries. To address these challenges, we propose a cross-modal knowledge distillation framework that leverages both text-to-text and speech-to-text channels to transfer knowledge from a text-based teacher model to a speech LLM. Extensive experiments on dialogue and audio understanding tasks validate the effectiveness of our approach in preserving textual knowledge, improving cross-modal alignment, and enhancing reasoning in speech-based interactions.

语音大模型知识蒸馏跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。