构建400亿词音乐语料库,提升语言模型在音乐领域的理解能力。
MuCPT: Music-related Natural Language Model Continued Pretraining
- 用轻量分类器筛选音乐相关文本,多阶段清洗并融合元数据。
- 通过参考模型软评分优化训练,减少噪声梯度,增强任务对齐。
- 提出MusicSimpleQA评测基准,适合音乐领域大模型评估与开发。
大型语言模型在通用任务中表现优异,但在音乐等专业领域仍受限,尤其在音乐娱乐场景下,语料规模、纯净度及数据与目标的匹配度至关重要。本文构建了一个包含400亿词的音乐相关自然语言语料库,融合开源与自有数据,并采用领域优先的数据处理流程:轻量级分类器过滤并加权域内文本,随后进行多阶段清洗、去重和隐私保护掩码。进一步整合多源音乐文本及其关联元数据,形成更广更结构化的领域知识基础。训练方面,引入基于参考模型(RM)的词级别软评分机制,统一损失比准则用于数据筛选与优化过程中的动态降权,降低噪声梯度,强化任务对齐信号,从而实现更高效的音乐领域持续预训练与对齐。为评估事实性,设计了MusicSimpleQA评测基准,采用短单答提示与自动一致率评分。除基准设计外,还系统比较了不同数据构成的影响。本工作同时推进了高质量语料与合理训练目标,提供可扩展的数据-训练框架与可复用的评估工具,助力音乐领域大模型建设。
原文摘要 · Abstract (English)
Large language models perform strongly on general tasks but remain constrained in specialized settings such as music, particularly in the music-entertainment domain, where corpus scale, purity, and the match between data and training objectives are critical. We address this by constructing a large, music-related natural language corpus (40B tokens) that combines open source and in-house data, and by implementing a domain-first data pipeline: a lightweight classifier filters and weights in-domain text, followed by multi-stage cleaning, de-duplication, and privacy-preserving masking. We further integrate multi-source music text with associated metadata to form a broader, better-structured foundation of domain knowledge. On the training side, we introduce reference-model (RM)-based token-level soft scoring for quality control: a unified loss-ratio criterion is used both for data selection and for dynamic down-weighting during optimization, reducing noise gradients and amplifying task-aligned signals, thereby enabling more effective music-domain continued pretraining and alignment. To assess factuality, we design the MusicSimpleQA benchmark, which adopts short, single-answer prompts with automated agreement scoring. Beyond the benchmark design, we conduct systematic comparisons along the axes of data composition. Overall, this work advances both the right corpus and the right objective, offering a scalable data-training framework and a reusable evaluation tool for building domain LLMs in the music field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。