arXiv:2508.06516cs.SDeess.AS2025-08

自动生成音乐混搭,用分离音轨判断搭配是否和谐。

AutoMashup: Automatic Music Mashups Creation

  • 通过分离音轨并评估其兼容性,自动生成混搭
  • 发现混搭兼容性具有不对称性,声乐与伴奏角色影响结果
  • 现有通用音频模型无法准确预测听感上的协调性

我们提出 AutoMashup,一个基于音轨分离、音乐分析和兼容性估计的自动混搭生成系统。采用 COCOLA 评估分离音轨间的兼容性,并探究通用预训练音频模型(CLAP 与 MERT)能否实现零样本的曲目对兼容性估计。结果表明,混搭兼容性具有不对称性——取决于音轨的角色(人声或伴奏);同时,当前嵌入表示无法复现 COCOLA 所测量的听觉连贯性。这些发现揭示了通用音频表征在混搭生成中进行兼容性估计的局限性。

原文摘要 · Abstract (English)

We introduce AutoMashup, a system for automatic mashup creation based on source separation, music analysis, and compatibility estimation. We propose using COCOLA to assess compatibility between separated stems and investigate whether general-purpose pretrained audio models (CLAP and MERT) can support zero-shot estimation of track pair compatibility. Our results show that mashup compatibility is asymmetric -- it depends on the role assigned to each track (vocals or accompaniment) -- and that current embeddings fail to reproduce the perceptual coherence measured by COCOLA. These findings underline the limitations of general-purpose audio representations for compatibility estimation in mashup creation.

音乐生成音频分离兼容性估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。