首个多音频评测基准,推动语音大模型处理多音源场景
Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models
- 构建20个数据集的多音频评测基准,覆盖11类真实场景任务
- 现有语音大模型在多音源下性能下降,新模型显著提升表现
- 用合成数据实现高数据效率,无需人工标注,适合实际部署
近期多种音频大语言模型(ALLMs)被提出,旨在通过统一模型同时处理各类音频任务。然而,现有评估主要聚焦单音频任务,而真实应用常需同时处理多个音频流。为填补这一差距,我们提出首个多音频评估(MAE)基准,包含来自11类多音频任务的20个数据集,涵盖语音与声音场景。在该基准上的综合实验表明,现有ALLMs虽能理解单个音频中的核心内容,但在多音频场景中表现不佳。为此,我们提出新型多音频-大语言模型(MALLM),利用自建合成数据进行判别式学习,有效捕捉多个相似音频间的上下文关系。结果表明,所提MALLM超越所有基线,且仅使用合成数据即实现高数据效率,无需人工标注。该模型为音频大模型迈向多音频处理时代铺路,助力机器更接近人类听觉能力。
原文摘要 · Abstract (English)
Various audio-LLMs (ALLMs) have been explored recently for tackling different audio tasks simultaneously using a single, unified model. While existing evaluations of ALLMs primarily focus on single-audio tasks, real-world applications often involve processing multiple audio streams simultaneously. To bridge this gap, we propose the first multi-audio evaluation (MAE) benchmark that consists of 20 datasets from 11 multi-audio tasks encompassing both speech and sound scenarios. Comprehensive experiments on MAE demonstrate that the existing ALLMs, while being powerful in comprehending primary audio elements in individual audio inputs, struggling to handle multi-audio scenarios. To this end, we propose a novel multi-audio-LLM (MALLM) to capture audio context among multiple similar audios using discriminative learning on our proposed synthetic data. The results demonstrate that the proposed MALLM outperforms all baselines and achieves high data efficiency using synthetic data without requiring human annotations. The proposed MALLM opens the door for ALLMs towards multi-audio processing era and brings us closer to replicating human auditory capabilities in machines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。