arXiv:2509.19631eess.AScs.AI2025-09

用强化学习提升多模态大模型的语音摘要能力

Advancing Speech Summarization in Multi-modal LLMs with Reinforcement Learning

  • 分阶段强化学习训练,直接从语音生成摘要
  • 性能超越多个更大模型,逼近顶尖文本模型
  • 适合研究语音理解与多模态AI的开发者

语音摘要是理解语音内容的关键环节,尤其在语音和音视频数据快速增长的背景下。近期基于多模态大语言模型(MLLM)的进展,利用大语言模型的能力,可直接从语音生成文本摘要,无需中间转录,同时支持可控风格和零样本泛化。然而,开源的MLLM仍显著落后于最先进的文本型大模型,限制了其在语音摘要中的实际应用。本文提出一种新颖的多阶段强化学习训练框架,显著增强MLLM的语音摘要能力。实验表明,该模型在多项指标上优于强基准,性能超过多个更大规模的MLLM,且大幅缩小了与顶尖文本型大模型的差距。

原文摘要 · Abstract (English)

Speech summarization is a critical component of spoken content understanding, particularly in the era of rapidly growing spoken and audiovisual data. Recent advances in multi-modal large language models (MLLMs), leveraging the power of LLMs, enable generating textual summaries directly from speech without intermediate transcriptions, while supporting controllable styles and zero-shot generalization. However, open-source MLLMs continue to lag behind the state-of-the-art text-based LLMs, limiting their practical deployment for speech summarization. In this work, we present a novel multi-stage reinforcement learning training framework to enhance the speech summarization capabilities in MLLMs. Our model delivers substantial improvements over strong baselines, outperforms much larger MLLMs, and significantly narrows the gap with state-of-the-art text-based LLMs.

语音摘要多模态强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。