arXiv:2603.24596eess.AScs.AI2026-03中稿 · Interspeech 2026被引 9

用文本模型指导语音大模型,显著缩小性能差距。

X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs

  • 让语音模型自探索,文本模型逐词反馈指导。
  • 复杂任务上性能接近文本模型,且保留原能力。
  • 适合研究语音大模型对齐与多模态训练的人。

从级联对话系统转向端到端(E2E)语音大语言模型(Speech LLMs)虽提升了延迟和副语言建模能力,但其性能常显著落后于文本模型。标准的监督微调(SFT)和强化学习(RL)方法难以弥合这一差距。为此,我们提出X-OPD,一种跨模态在线策略蒸馏框架,旨在系统性对齐语音模型与文本模型的能力。X-OPD通过在策略回溯中让语音模型自主探索,由文本教师模型评估轨迹并提供逐标记反馈,有效将教师模型的能力蒸馏至学生模型的多模态表征中。多基准实验表明,X-OPD显著缩小了复杂任务中的性能差距,同时保持了模型的固有能力。

原文摘要 · Abstract (English)

While the shift from cascaded dialogue systems to end-to-end (E2E) speech Large Language Models (LLMs) improves latency and paralinguistic modeling, E2E models often exhibit a significant performance degradation compared to their text-based counterparts. The standard Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) training methods fail to close this gap. To address this, we propose X-OPD, a novel Cross-Modal On-Policy Distillation framework designed to systematically align the capabilities of Speech LLMs to their text-based counterparts. X-OPD enables the Speech LLM to explore its own distribution via on-policy rollouts, where a text-based teacher model evaluates these trajectories and provides token-level feedback, effectively distilling teacher's capabilities into student's multi-modal representations. Extensive experiments across multiple benchmarks demonstrate that X-OPD significantly narrows the gap in complex tasks while preserving the model's inherent capabilities.

语音大模型跨模态蒸馏能力对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。