arXiv:2501.16171eess.AScs.IR2025-01被引 2

用超椭球查询实现任意音乐源分离,突破传统固定音轨限制。

Separate This, and All of these Things Around It: Music Source Separation via Hyperellipsoidal Queries

  • 以超椭球区域作为查询,灵活指定目标音源的位置与范围。
  • 在MoisesDB上达到当前最佳的信噪比与检索指标表现。
  • 适合需要自定义分离目标的音乐制作与音频编辑场景。

音乐源分离是将混合音乐音频中一个或多个组成部分(即'声部')提取出来的任务。传统方法依赖固定声部范式,主流系统仅支持人声、鼓、贝斯和'其他'这四类声部。近年来,研究开始挑战这一范式,提出可通过额外查询输入来指定目标声音类型,实现任意声音的分离。本文进一步提出基于'按区域查询'的分离系统,通过超椭球区域作为查询,可直观且参数化地定义目标位置及其扩展范围,无需预设声部类别。在MoisesDB数据集上的评估表明,该方法在信噪比与检索指标上均达到当前最优性能。

原文摘要 · Abstract (English)

Music source separation is an audio-to-audio retrieval task of extracting one or more constituent components, or composites thereof, from a musical audio mixture. Each of these constituent components is often referred to as a "stem" in literature. Historically, music source separation has been dominated by a stem-based paradigm, leading to most state-of-the-art systems being either a collection of single-stem extraction models, or a tightly coupled system with a fixed, difficult-to-modify, set of supported stems. Combined with the limited data availability, advances in music source separation have thus been mostly limited to the "VDBO" set of stems: \textit{vocals}, \textit{drum}, \textit{bass}, and the catch-all \textit{others}. Recent work in music source separation has begun to challenge the fixed-stem paradigm, moving towards models able to extract any musical sound as long as this target type of sound could be specified to the model as an additional query input. We generalize this idea to a \textit{query-by-region} source separation system, specifying the target based on the query regardless of how many sound sources or which sound classes are contained within it. To do so, we propose the use of hyperellipsoidal regions as queries to allow for an intuitive yet easily parametrizable approach to specifying both the target (location) as well as its spread. Evaluation of the proposed system on the MoisesDB dataset demonstrated state-of-the-art performance of the proposed system both in terms of signal-to-noise ratios and retrieval metrics.

音乐分离查询分割超椭球音频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。