arXiv:2507.14815cs.CL2025-07NeurIPS被引 3

让大模型高效处理长语音,无需额外训练数据。

FastLongSpeech: Enhancing Large Speech-Language Models for Efficient Long-Speech Processing

  • 通过迭代融合压缩长语音序列,降低计算负担。
  • 动态压缩训练使模型在不同压缩比下保持性能,提升泛化能力。
  • 适合需要高效长语音理解的场景,如会议转录、播客分析。

大型语言模型(LLMs)的快速发展推动了大型语音-语言模型(LSLMs)的进步,提升了语音理解和生成能力。尽管现有LSLMs多聚焦于语音生成或短语音任务,但高效处理长语音仍是关键挑战,主要受限于长语音训练数据稀缺和长序列带来的高计算成本。为此,我们提出FastLongSpeech框架,无需专用长语音训练数据即可扩展LSLM对长语音的处理能力。该框架采用迭代融合策略,将过长的语音序列压缩至可管理长度;并通过动态压缩训练,让模型在不同压缩比的短语音序列中学习,实现从短语音到长语音任务的能力迁移。为评估模型在长语音上的表现,我们构建了名为LongSpeech-Eval的评测基准。实验表明,该方法在长语音与短语音任务中均表现优异,且显著提升推理效率。

原文摘要 · Abstract (English)

The rapid advancement of Large Language Models (LLMs) has spurred significant progress in Large Speech-Language Models (LSLMs), enhancing their capabilities in both speech understanding and generation. While existing LSLMs often concentrate on augmenting speech generation or tackling a diverse array of short-speech tasks, the efficient processing of long-form speech remains a critical yet underexplored challenge. This gap is primarily attributed to the scarcity of long-speech training datasets and the high computational costs associated with long sequences. To address these limitations, we introduce FastLongSpeech, a novel framework designed to extend LSLM capabilities for efficient long-speech processing without necessitating dedicated long-speech training data. FastLongSpeech incorporates an iterative fusion strategy that can compress excessively long-speech sequences into manageable lengths. To adapt LSLMs for long-speech inputs, it introduces a dynamic compression training approach, which exposes the model to short-speech sequences at varying compression ratios, thereby transferring the capabilities of LSLMs to long-speech tasks. To assess the long-speech capabilities of LSLMs, we develop a long-speech understanding benchmark called LongSpeech-Eval. Experiments show that our method exhibits strong performance in both long-speech and short-speech tasks, while greatly improving inference efficiency.

语音处理大模型长语音高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。