arXiv:2609.02343cs.SDcs.CL2026-09

构建1500万条细粒度音频描述,提升音频检索效果

SonicCaps: Large-Scale Diverse and Fine-Grained Captioning for Improved Audio-Retrieval

论文配图:SonicCaps: Large-Scale Diverse and Fine-Grained Captioning for Improved Audio-Retrieval
图 1 · 摘自论文原文
  • 用多模态大模型生成每段音频24条不同风格的描述
  • 人类评估显示新数据更精准细致,显著优于现有数据集
  • 基于此训练的CLAP模型在多个任务中表现更强

近年来,音频-语言建模的进步依赖于大规模音频描述数据集。然而,现有数据集仍受限于语义多样性低、描述泛化、且音频与描述为一对一映射,无法反映听觉感知的固有模糊性。我们提出SonicCaps,一个包含约1500万条描述和约70万段音频的大型音频描述数据集,使用多模态大语言模型(Qwen3-Omni)在音频和文本共同条件下生成。为促进多样性,通过结构化提示工程和少样本生成,每段音频生成约24条描述,涵盖主描述、改写变体(冗余度、风格)和语义标签。人工评估表明,SonicCaps在描述性与精确性上均显著优于现有数据集,且与质量评分高度相关。最终,使用多描述采样策略在SonicCaps上训练的CLAP模型,在音频检索和零样本分类任务中持续表现更优,并在公开及商业基准上展现出更强泛化能力。我们已在Hugging Face发布SonicCaps及两个专用CLAP模型:https://huggingface.co/datasets/Zineb/SonicCaps。

原文摘要 · Abstract (English)

Recent advances in audio-language modeling have been driven by large-scale audio captioning datasets. However, existing datasets remain limited by low semantic diversity, generic descriptions lacking acoustic details, and one-to-one audio-caption mappings that poorly reflect the inherent ambiguity of auditory perception. We introduce SonicCaps, a large-scale audio captioning dataset comprising ~15M captions paired with ~700k audio clips, generated using a multi-modal large language model (Qwen3-Omni) conditioned on both audio and text. To explicitly promote diversity, we generate around 24 captions per audio via structured prompt engineering and few- shot generation, spanning main descriptions, rephrased variants (verbosity, style) and semantic tags. Human evaluation shows that SonicCaps is rated significantly higher than existing captioning datasets, with fine-grained analyses indicating that our captions are perceived as more descriptive and precise, which strongly correlates with quality judgments. Finally, training CLAP models on SonicCaps with a multi-caption sampling strategy consistently improves audio retrieval and zero-shot classification, with stronger generalization across public and commercial benchmarks. We release both SonicCaps and two specialized CLAP models on hugging face: https://huggingface.co/datasets/Zineb/SonicCaps.

音频描述多模态数据集检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。