arXiv:2412.12009eess.AScs.AI2024-12中稿 · IEEE ICME 2025被引 22

提出语音感知的令牌剪枝方法,提升长音频理解效率

SpeechPrune: Context-aware Token Pruning for Speech Information Retrieval

  • 基于语音-文本相似度与注意力近似值,无训练剪枝冗余音频令牌
  • 在90秒音频上,剪枝20%时准确率提升29%,最高达47%
  • 支持80%剪枝率下仍保持性能,适合资源受限的长语音场景

我们引入语音信息检索(SIR)这一新任务,评估语音大模型在长上下文下的表现,并构建了包含1,012个样本的SPIRAL基准,用于测试模型从约90秒语音输入中提取关键信息的能力。当前语音大模型在短任务中表现优异,但在处理长音频序列时面临计算与表征挑战。为此,我们提出SpeechPrune,一种无需训练的令牌剪枝策略,通过语音-文本相似度和近似注意力分数,高效剔除无关令牌。在SPIRAL基准上,SpeechPrune在20%剪枝率下相较原模型和随机剪枝模型分别实现29%和47%的准确率提升;即使在80%剪枝率下,仍能维持网络性能。该方法展示了令牌级剪枝在高效、可扩展的长语音理解中的潜力。

原文摘要 · Abstract (English)

We introduce Speech Information Retrieval (SIR), a new long-context task for Speech Large Language Models (Speech LLMs), and present SPIRAL, a 1,012-sample benchmark testing models' ability to extract critical details from approximately 90-second spoken inputs. While current Speech LLMs excel at short-form tasks, they struggle with the computational and representational demands of longer audio sequences. To address this limitation, we propose SpeechPrune, a training-free token pruning strategy that uses speech-text similarity and approximated attention scores to efficiently discard irrelevant tokens. In SPIRAL, SpeechPrune achieves accuracy improvements of 29% and up to 47% over the original model and the random pruning model at a pruning rate of 20%, respectively. SpeechPrune can maintain network performance even at a pruning level of 80%. This approach highlights the potential of token-level pruning for efficient and scalable long-form speech understanding.

语音理解令牌剪枝长序列高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。