用多模态大模型自动发现可解释的音频特征,提升低资源分类效率。
Adaptive Discovery of Interpretable Audio Attributes with Multimodal LLMs for Low-Resource Classification
- 用多模态大模型替代人工,动态提示识别关键声学特征
- 在多个音频任务中优于直接大模型预测,训练仅需11分钟
- 适合需要可解释性的低资源音频分类场景
在低资源音频分类的预测建模中,提取高精度且可解释的音频属性至关重要,尤其在高可靠性应用中不可或缺。尽管人工驱动的属性发现有效,但其处理速度慢成为瓶颈。本文提出一种基于多模态大语言模型(MLLMs)的自适应可解释音频属性发现方法。通过将人类替换为MLLMs,该方法显著加快属性发现速度。利用提示工程动态识别显著声学特征,并构建基于属性的集成分类器。在多个音频任务上的实验结果表明,该方法在多数情况下优于直接使用MLLMs进行预测。整个训练过程仅耗时11分钟,证明其是一种高效、自适应的解决方案,超越传统依赖人工的方法。
原文摘要 · Abstract (English)
In predictive modeling for low-resource audio classification, extracting high-accuracy and interpretable attributes is critical. Particularly in high-reliability applications, interpretable audio attributes are indispensable. While human-driven attribute discovery is effective, its low throughput becomes a bottleneck. We propose a method for adaptively discovering interpretable audio attributes using Multimodal Large Language Models (MLLMs). By replacing humans in the AdaFlock framework with MLLMs, our method achieves significantly faster attribute discovery. Our method dynamically identifies salient acoustic characteristics via prompting and constructs an attribute-based ensemble classifier. Experimental results across various audio tasks demonstrate that our method outperforms direct MLLM prediction in the majority of evaluated cases. The entire training completes within 11 minutes, proving it a practical, adaptive solution that surpasses conventional human-reliant approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。