用自生成对齐策略打造通用音频大模型,兼顾语言能力与听觉理解。
DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment
- 让大语言模型自动生成音频任务标签,避免遗忘原有语言能力。
- 构建500万样本跨模态数据集,覆盖50个音源类型,时长超7000小时。
- 无需微调即可零样本应对多种音频任务,适合多场景通用音频系统。
我们提出DeSTA2.5-Audio,一种面向鲁棒听觉感知与指令遵循的通用大型音频语言模型(LALM)。现有LALM在引入音频能力时常导致大语言模型(LLM)原有能力出现灾难性遗忘。为此,我们重构数据构建流程,提出自生成跨模态对齐策略——由骨干LLM自动生成训练目标,命名为DeSTA,旨在保留其原生语言能力,实现无需特定任务微调的零样本泛化。我们构建了DeSTA-AQA5M数据集,包含500万条样本,源自7000小时音频,涵盖50种不同数据集,包括语音、环境声和音乐。DeSTA2.5-Audio在Dynamic-SUPERB、MMAU、SAKURA、Speech-IFEval和VoiceBench等多个音频-语言基准上达到领先或竞争力表现。对比实验证明,自生成策略优于现有方法。研究强调了精心设计数据构造在LALM开发中的重要性,并为构建稳健的通用型音频大模型提供实用洞见。
原文摘要 · Abstract (English)
We introduce DeSTA2.5-Audio, a general-purpose Large Audio Language Model (LALM) designed for robust auditory perception and instruction-following. Recent LALMs augment Large Language Models (LLMs) with auditory capabilities by training on large-scale audio-instruction datasets. However, existing LALMs have often suffered from the catastrophic forgetting of the LLM's original abilities. Therefore, balancing knowledge retention and audio perception has become a critical challenge. To address this, we revisit the data construction pipeline and propose a self-generated cross-modal alignment strategy in which the backbone LLM generates its own training targets, named DeSTA. This approach aims at preserving the LLM's native language proficiency thereby enabling zero-shot generalization without task-specific tuning. We construct DeSTA-AQA5M, a large-scale, task-agnostic dataset containing 5 million training samples derived from 7,000 hours of audio spanning 50 diverse datasets, including speech, environmental sounds, and music. DeSTA2.5-Audio achieves state-of-the-art or competitive performance across a wide range of audio-language benchmarks, including Dynamic-SUPERB, MMAU, SAKURA, Speech-IFEval, and VoiceBench. Comprehensive comparative studies demonstrate that our self-generated strategy outperforms existing training strategies. Our findings underscore the importance of carefully designed data construction in LALM development and offer practical insights for building robust, general-purpose LALMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。