arXiv:2603.24038eess.AScs.SD2026-03中稿 · ICASSP 2026被引 5

构建大规模细粒度音频描述数据集,提升音频理解模型泛化能力

ACAVCaps: Enabling large-scale training for fine-grained and diverse audio understanding

  • 多专家管道分析音频的语音、音乐与声学特征,生成丰富描述
  • 基于该数据集预训练的模型在多个下游任务中表现显著更优
  • 适合研究音频语言模型、细粒度音频理解的开发者使用

通用音频理解是大型音频-语言模型的核心目标,音频字幕任务是其发展的基石。然而,当前数据集规模不足且描述粒度不够,制约了模型发展。为此,我们提出ACAVCaps,一个大规模、细粒度、多维度的音频字幕数据集。该数据集源自ACAV100M,通过多专家管道从语音、音乐和声学属性等多角度分析音频,并由大语言模型合成丰富详尽的描述。实验表明,基于ACAVCaps预训练的模型在多个下游任务上展现出显著更强的泛化能力,优于在其他主流字幕数据集上训练的模型。数据集已开源:https://github.com/xiaomi-research/acavcaps。

原文摘要 · Abstract (English)

General audio understanding is a fundamental goal for large audio-language models, with audio captioning serving as a cornerstone task for their development. However, progress in this domain is hindered by existing datasets, which lack the scale and descriptive granularity required to train truly versatile models. To address this gap, we introduce ACAVCaps, a new large-scale, fine-grained, and multi-faceted audio captioning dataset. Derived from the ACAV100M collection, ACAVCaps is constructed using a multi-expert pipeline that analyzes audio from diverse perspectives-including speech, music, and acoustic properties-which are then synthesized into rich, detailed descriptions by a large language model. Experimental results demonstrate that models pre-trained on ACAVCaps exhibit substantially stronger generalization capabilities on various downstream tasks compared to those trained on other leading captioning datasets. The dataset is available at https://github.com/xiaomi-research/acavcaps.

音频理解数据集多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。