用人类与多模态大模型协作自动发现音频场景标签,降低标注成本。
AuditoryHuM: Auditory Scene Label Generation and Clustering using Human-MLLM Collaboration
- 通过大模型生成音频标签,结合人类反馈优化质量
- 在三个数据集上实现可复用的标准化场景分类体系
- 适合需要轻量级音频识别的边缘设备应用
人工标注音频数据耗时费力,且难以平衡标签粒度与声学可分性。我们提出AuditoryHuM,一种基于人类-多模态大语言模型(MLLM)协作的无监督音频场景标签发现与聚类框架。利用Gemma和Qwen等MLLM生成上下文相关的音频标签,通过零样本学习方法(Human-CLAP)量化生成文本标签与原始音频内容的对齐程度,针对对齐度最低的样本进行针对性的人工干预以提升标签质量。采用改进的轮廓系数(引入惩罚参数)对发现的标签进行主题一致性聚类,以平衡簇内凝聚度与主题粒度。在ADVANCE、AHEAD-DS和TAU 2019三个多样化音频场景数据集上验证,该框架提供了可扩展、低成本的标准化分类体系,支持训练轻量级场景识别模型,适用于助听器、智能家居助手等边缘设备部署。
原文摘要 · Abstract (English)
Manual annotation of audio datasets is labour intensive, and it is challenging to balance label granularity with acoustic separability. We introduce AuditoryHuM, a novel framework for the unsupervised discovery and clustering of auditory scene labels using a collaborative Human-Multimodal Large Language Model (MLLM) approach. By leveraging MLLMs (Gemma and Qwen) the framework generates contextually relevant labels for audio data. To ensure label quality and mitigate hallucinations, we employ zero-shot learning techniques (Human-CLAP) to quantify the alignment between generated text labels and raw audio content. A strategically targeted human-in-the-loop intervention is then used to refine the least aligned pairs. The discovered labels are grouped into thematically cohesive clusters using an adjusted silhouette score that incorporates a penalty parameter to balance cluster cohesion and thematic granularity. Evaluated across three diverse auditory scene datasets (ADVANCE, AHEAD-DS, and TAU 2019), AuditoryHuM provides a scalable, low-cost solution for creating standardised taxonomies. This solution facilitates the training of lightweight scene recognition models deployable to edge devices, such as hearing aids and smart home assistants. The project page and code: https://github.com/Australian-Future-Hearing-Initiative
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。