开源中文音乐数据集CCMusic,助力中文音乐信息检索研究。
CCMusic: An Open and Diverse Database for Chinese Music Information Retrieval Research
- 整合公开与未公开数据,统一格式并清洗标注。
- 提供分类与检测任务的标准化评测框架。
- 适合作为中文音乐分析、检索任务的研究起点。
数据在计算机相关领域至关重要,音乐信息检索(MIR)作为计算机科学与音乐的交叉领域尤为依赖高质量数据。本文介绍CCMusic,一个面向中文音乐任务的开放且多样化的数据库,涵盖多个专为中文音乐设计的数据集。该数据库融合了已发表与未发表的数据,通过数据清洗、标签优化和结构统一,确保数据一致性并生成即用版本。我们为所有数据集构建了统一的评估框架,支持分类与检测任务,实现跨数据集的标准性与可复现性结果。数据库托管于HuggingFace和ModelScope两个开放平台,便于获取与使用。
原文摘要 · Abstract (English)
Data are crucial in various computer-related fields, including music information retrieval (MIR), an interdisciplinary area bridging computer science and music. This paper introduces CCMusic, an open and diverse database comprising multiple datasets specifically designed for tasks related to Chinese music, highlighting our focus on this culturally rich domain. The database integrates both published and unpublished datasets, with steps taken such as data cleaning, label refinement, and data structure unification to ensure data consistency and create ready-to-use versions. We conduct benchmark evaluations for all datasets using a unified evaluation framework developed specifically for this purpose. This publicly available framework supports both classification and detection tasks, ensuring standardized and reproducible results across all datasets. The database is hosted on HuggingFace and ModelScope, two open and multifunctional data and model hosting platforms, ensuring ease of accessibility and usability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。