arXiv:2505.20166eess.AScs.AI2025-05被引 7

用合成数据提升音频语言模型的对齐能力,减少幻觉并保持文本理解力。

From Alignment to Advancement: Bootstrapping Audio-Language Alignment with Synthetic Data

  • 利用大模型生成对比性合成数据,增强模型区分真实与虚假声音的能力。
  • 在多个基准上有效降低音频幻觉率,同时保持指令遵循和音频理解性能。
  • 支持多音频输入场景,适合需要跨模态推理的研究者与开发者。

音频感知的大语言模型(ALLMs)在理解和处理音频输入方面取得了显著进展。这些模型通常通过在音频相关任务上额外训练,从文本型大语言模型(LLMs)演化而来。然而,这一过程存在两大局限:其一,ALLMs常出现灾难性遗忘,导致关键文本能力如指令遵循能力下降,甚至在输入中不存在时产生幻觉声音,影响可靠性;其二,实现音视频跨模态对齐通常依赖大量特定任务的问答对进行指令微调,资源消耗大。此前工作曾利用骨干LLM生成通用、描述性的对齐数据。本文提出一种数据生成框架,可生成类对比学习的训练数据,旨在提升ALLMs区分实际存在与缺失声音的能力。进一步扩展至多音频场景,使模型能够解释音频间的差异或生成统一描述所有输入的标题,从而增强音视频对齐。整个训练框架称为基于骨干LLM合成数据的自举式音视频对齐(BALSa)。实验表明,该方法能有效缓解音频幻觉,同时在音频理解与推理基准及指令遵循任务上保持强性能。引入多音频训练还进一步提升了模型的综合理解与推理能力。总体而言,BALSa为构建高效、可扩展的ALLMs提供了一种新路径。

原文摘要 · Abstract (English)

Audio-aware large language models (ALLMs) have recently made great strides in understanding and processing audio inputs. These models are typically adapted from text-based large language models (LLMs) through additional training on audio-related tasks. This adaptation process presents two major limitations. First, ALLMs often suffer from catastrophic forgetting, where crucial textual capabilities like instruction-following are lost after training on audio data. In some cases, models may even hallucinate sounds that are not present in the input audio, raising concerns about reliability. Second, achieving cross-modal alignment between audio and language typically relies on large collections of task-specific question-answer pairs for instruction tuning, making it resource-intensive. To address these issues, previous works have leveraged the backbone LLMs to synthesize general-purpose, caption-style alignment data. In this paper, we propose a data generation framework that produces contrastive-like training data, designed to enhance ALLMs' ability to differentiate between present and absent sounds. We further extend our approach to multi-audio scenarios, enabling the model to either explain differences between audio inputs or produce unified captions that describe all inputs, thereby enhancing audio-language alignment. We refer to the entire ALLM training framework as bootstrapping audio-language alignment via synthetic data generation from backbone LLMs (BALSa). Experimental results indicate that our method effectively mitigates audio hallucinations while reliably maintaining strong performance on audio understanding and reasoning benchmarks, as well as instruction-following skills. Moreover, incorporating multi-audio training further enhances the model's comprehension and reasoning capabilities. Overall, BALSa offers an efficient and scalable approach to developing ALLMs.

音视频对齐合成数据大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。