用结构化元数据提升音效设计中的音频语言模型性能
Audiocards: Structured Metadata Improves Audio Language Models For Sound Design
- 提出 audiocards,基于声学属性的结构化元数据
- 在专业音效库上提升检索与描述生成效果
- 适合音效设计、音频标注与多模态研究者
音效设计师在大型音效库中搜索声音时,常依赖音效类别或视觉情境等元数据。但这些元数据往往缺失或不完整,需大量人工添加。现有自动化方案(如自动生成描述、基于嵌入的文本-音频检索)未针对音效设计需求训练。为此,我们提出 audiocards——利用大语言模型的世界知识,构建基于声学属性与声学描述符的结构化元数据。实验证明,使用 audiocards 训练可显著提升下游任务表现:包括文本-音频检索、描述性标题生成及元数据生成,在专业音效库上的性能优于基线单句标题方法。此外,该方法也提升了通用音频描述与检索能力。我们发布了经过精心整理的音效 audiocards 数据集,以推动音效设计中音频语言建模的研究。
原文摘要 · Abstract (English)
Sound designers search for sounds in large sound effects libraries using aspects such as sound class or visual context. However, the metadata needed for such search is often missing or incomplete, and requires significant manual effort to add. Existing solutions to automate this task by generating metadata, i.e. captioning, and search using learned embeddings, i.e. text-audio retrieval, are not trained on metadata with the structure and information pertinent to sound design. To this end we propose audiocards, structured metadata grounded in acoustic attributes and sonic descriptors, by exploiting the world knowledge of LLMs. We show that training on audiocards improves downstream text-audio retrieval, descriptive captioning, and metadata generation on professional sound effects libraries. Moreover, audiocards also improve performance on general audio captioning and retrieval over the baseline single-sentence captioning approach. We release a curated dataset of sound effects audiocards to invite further research in audio language modeling for sound design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。