构建新任务与数据集,测试大模型对音频与语音的联合推理能力
What Are They Doing? Joint Audio-Speech Co-Reasoning
- 提出联合音视频共推理任务JASCO,要求模型同步分析声音与人声
- 构建包含场景理解的'What Are They Doing'数据集,支持多模态推理评测
- 揭示大模型对音频和语音模态的依赖关系,为多模态研究提供洞见
在音频与语音处理中,任务通常仅关注音频或语音模态,即使同一音频片段中同时存在声音与人类语音。近年来,听觉大语言模型(ALLMs)使得在同一模型中同时处理音频与语音成为可能,从而引发了对联合音频-语音任务的进一步探索。本文建立了一个新基准,用于研究ALLMs在联合音频-语音处理中的表现。具体而言,我们提出了联合音频-语音共推理(JASCO)任务,该任务统一了音频与语音处理,严格要求跨两个模态的协同推理。我们还发布了名为“What Are They Doing”的场景推理数据集。此外,通过分析模型对各模态的依赖性,提供了对模型行为的深入洞察。
原文摘要 · Abstract (English)
In audio and speech processing, tasks usually focus on either the audio or speech modality, even when both sounds and human speech are present in the same audio clip. Recent Auditory Large Language Models (ALLMs) have made it possible to process audio and speech simultaneously within a single model, leading to further considerations of joint audio-speech tasks. In this paper, we establish a novel benchmark to investigate how well ALLMs can perform joint audio-speech processing. Specifically, we introduce Joint Audio-Speech Co-Reasoning (JASCO), a novel task that unifies audio and speech processing, strictly requiring co-reasoning across both modalities. We also release a scene-reasoning dataset called "What Are They Doing". Additionally, we provide deeper insights into the models' behaviors by analyzing their dependence on each modality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。