arXiv:2503.15338eess.AScs.CL2025-03被引 2

让大模型同时听懂语音指令和环境声音,提升真实场景下的理解能力。

Solla: Towards a Speech-Oriented LLM That Hears Acoustic Context

  • 融合音频事件识别与语音助手预测,实现语音与音频同步理解。
  • 在难易双难度测试集上表现优于或持平基线模型。
  • 适合需要语音+环境音协同分析的应用场景,如智能助手、无障碍系统。

大型语言模型(LLMs)近年来展现出处理文本及多模态输入(如语音和音频)的惊人能力。然而,现有大多数模型主要依赖文本指令分析输入信号,忽视了语音指令与音频混合输入的现实场景。为解决这一问题,我们提出Solla,一种新型框架,可同时理解基于语音的问题并感知声学上下文。Solla引入音频标记模块以有效识别和表征音频事件,并采用语音识别辅助的预测方法,提升对口语内容的理解。为严格评估Solla及其他公开模型,我们构建了新基准数据集SA-Eval,包含三项任务:音频事件分类、音频描述生成和音频问答。SA-Eval涵盖多种说话风格的语音指令,分为易、难两个难度层级,以覆盖真实声学条件。实验结果表明,Solla在易与难测试集上的表现均达到或超过基线模型,验证了其在联合理解语音与音频方面的有效性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have recently shown remarkable ability to process not only text but also multimodal inputs such as speech and audio. However, most existing models primarily focus on analyzing input signals using text instructions, overlooking scenarios in which speech instructions and audio are mixed and serve as inputs to the model. To address these challenges, we introduce Solla, a novel framework designed to understand speech-based questions and hear the acoustic context concurrently. Solla incorporates an audio tagging module to effectively identify and represent audio events, as well as an ASR-assisted prediction method to improve comprehension of spoken content. To rigorously evaluate Solla and other publicly available models, we propose a new benchmark dataset called SA-Eval, which includes three tasks: audio event classification, audio captioning, and audio question answering. SA-Eval has diverse speech instruction with various speaking styles, encompassing two difficulty levels, easy and hard, to capture the range of real-world acoustic conditions. Experimental results show that Solla performs on par with or outperforms baseline models on both the easy and hard test sets, underscoring its effectiveness in jointly understanding speech and audio.

语音理解多模态听觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。