arXiv:2506.01111cs.SDcs.AI2025-06被引 10

用多模态融合提升音频描述细节,生成更精准的音频解说。

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion

  • 分两阶段:先提取语音、音乐、视觉等多模态信息,再由大模型整合生成描述。
  • 构建120万条精细音频描述数据集,含600万问答对,规模领先。
  • 适合做音频理解、内容检索与多模态交互的研究者使用。

高质量的大规模音频描述对推动音频理解至关重要,但现有自动化方法常因依赖有限的单模态或浅层多模态信息,导致生成的描述缺乏细粒度和上下文准确性。受人类听觉感知启发,我们提出一种新颖的两阶段自动化流程:首先利用专用预训练模型提取多样化的上下文线索(如语音、音乐、通用声音及关联视频的视觉信息),随后通过大型语言模型(LLM)融合这些丰富的多模态输入,生成详尽且具备上下文意识的音频描述。本工作主要贡献包括:(1)一种可扩展的细粒度音频描述生成方法;(2)FusionAudio,一个包含120万条详细描述与600万问答对的新大规模数据集;(3)基于CLAP的音频编码器,具有更强的音文对齐能力与指令遵循性能。该研究为复杂音频环境的更细致、准确的自动理解开辟了道路。代码与数据见 https://github.com/satsuki2486441738/FusionAudio。

原文摘要 · Abstract (English)

High-quality, large-scale audio captioning is crucial for advancing audio understanding, yet current automated methods often generate captions that lack fine-grained detail and contextual accuracy, primarily due to their reliance on limited unimodal or superficial multimodal information. Drawing inspiration from human auditory perception, which adeptly integrates cross-modal cues and performs sophisticated auditory scene analysis, we introduce a novel two-stage automated pipeline. This pipeline first employs specialized pretrained models to extract diverse contextual cues (e.g., speech, music, general sounds, and visual information from associated video). A large language model (LLM) then synthesizes these rich, multimodal inputs to generate detailed and context-aware audio captions. Key contributions of this work include: (1) the proposed scalable method for fine-grained audio caption generation; (2) FusionAudio, a new large-scale dataset comprising 1.2 million such detailed captions, combined with 6 million QA pairs; and (3) enhanced audio models developed using FusionAudio, specifically a CLAP-based audio encoder with superior audio-text alignment and instruction following. This paper paves the way for more nuanced and accurate automated understanding of complex audio environments. Code and data can be found in https://github.com/satsuki2486441738/FusionAudio.

音频描述多模态大模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。