arXiv:2603.04293cs.SDcs.AI2026-03中稿 · NLP4MusA 2026

LabelBuddy用AI辅助实现音乐音频的协同智能标注,解决主观标注难问题。

LabelBuddy: An Open Source Music and Audio Language Annotation Tagging Tool Using AI Assistance

  • 通过容器化后端分离界面与推理,支持自定义AI模型预标注。
  • 支持多用户协作共识,提升标注一致性与效率。
  • 开源工具适合音乐信息检索研究者与音频数据标注团队使用。

机器学习、大型音频语言模型(LALMs)和自主AI代理在音乐信息检索(MIR)中的发展,要求从静态标签转向丰富且符合人类意图的表征学习。然而,缺乏能捕捉音频标注主观细微差别的开源基础设施,仍是关键瓶颈。本文介绍 extbf{LabelBuddy},一个开源的协作式自动标注音频标注工具,旨在弥合人类意图与机器理解之间的差距。不同于静态工具,它通过容器化后端将界面与推理解耦,允许用户接入自定义模型进行AI辅助预标注。我们描述了系统架构,支持多用户共识、模型容器隔离,并规划了扩展代理和LALMs的路线图。代码已公开于 https://github.com/GiannisProkopiou/gsoc2022-Label-buddy。

原文摘要 · Abstract (English)

The advancement of Machine learning (ML), Large Audio Language Models (LALMs), and autonomous AI agents in Music Information Retrieval (MIR) necessitates a shift from static tagging to rich, human-aligned representation learning. However, the scarcity of open-source infrastructure capable of capturing the subjective nuances of audio annotation remains a critical bottleneck. This paper introduces \textbf{LabelBuddy}, an open-source collaborative auto-tagging audio annotation tool designed to bridge the gap between human intent and machine understanding. Unlike static tools, it decouples the interface from inference via containerized backends, allowing users to plug in custom models for AI-assisted pre-annotation. We describe the system architecture, which supports multi-user consensus, containerized model isolation, and a roadmap for extending agents and LALMs. Code available at https://github.com/GiannisProkopiou/gsoc2022-Label-buddy.

音频标注AI辅助开源工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。