arXiv:2505.18517cs.AIcs.LG2025-05被引 2

让大模型更懂语音,用可学习的软标记实现高效多任务音频理解。

LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs

  • 用可学习的键值对动态选择提示,平衡通用与任务特定知识。
  • 仅需少量参数即可达到竞争力性能,训练过程简化为单阶段。
  • 提升模型可解释性,分析不同任务间提示的选择差异。

基于大规模语言模型(LLMs)的基础模型在处理多种任务和模态方面表现出色。然而,由于声学环境差异和任务变化,将其适配于通用语音-语言任务仍具挑战。本文提出LiSTEN(Learning Soft Token Embeddings for Neural Audio LLMs),一种将LLMs适配至语音与音频任务的框架。LiSTEN采用可学习的键值对动态提示选择策略,使模型在多任务设置中平衡通用知识与任务特定知识,同时避免过拟合。该方法减少对大规模语音识别或字幕数据集的依赖,以更少的可训练参数实现竞争性性能,并通过单阶段训练流程简化训练过程。此外,LiSTEN通过分析不同任务间所选提示的多样性与重叠度,增强了模型的可解释性。

原文摘要 · Abstract (English)

Foundation models based on large language models (LLMs) have shown great success in handling various tasks and modalities. However, adapting these models for general-purpose audio-language tasks is challenging due to differences in acoustic environments and task variations. In this work, we introduce LiSTEN Learning Soft Token Embeddings for Neural Audio LLMs), a framework for adapting LLMs to speech and audio tasks. LiSTEN uses a dynamic prompt selection strategy with learnable key-value pairs, allowing the model to balance general and task-specific knowledge while avoiding overfitting in a multitask setting. Our approach reduces dependence on large-scale ASR or captioning datasets, achieves competitive performance with fewer trainable parameters, and simplifies training by using a single-stage process. Additionally, LiSTEN enhances interpretability by analyzing the diversity and overlap of selected prompts across different tasks.

音频理解大模型软标记多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。