arXiv:2507.00155eess.AScs.SD2025-07被引 1

测试主流音乐分离模型在立体声与双耳音频中的空间信息保留能力

Do Music Source Separation Models Preserve Spatial Information in Binaural Audio?

  • 用MUSDB18-HQ和头相关传输函数合成双耳数据
  • 现有模型在双耳音频中空间信息严重丢失,依赖模型结构和乐器类型
  • 为沉浸式音频与无障碍应用提供新研究方向

双耳音频在音乐信息检索领域仍被忽视。鉴于虚拟现实与增强现实体验的兴起及无障碍应用潜力,我们研究了现有音乐源分离(MSS)模型在双耳音频上的表现。尽管这些模型处理双通道输入,但其对空间信息的保留效果尚不明确。本文评估了几种主流MSS模型在标准立体声与新型双耳数据集上的空间信息保持能力。双耳数据通过MUSDB18-HQ音轨与开源头相关传输函数合成,乐器源随机分布于水平平面。采用信号处理与双耳线索指标评估分离音轨的空间质量。结果表明,立体声MSS模型无法有效保留双耳音频所需的沉浸感关键空间信息,退化程度受模型架构及目标乐器影响。最后,我们指出了MSS与沉浸式音频交叉领域的潜在研究机遇。

原文摘要 · Abstract (English)

Binaural audio remains underexplored within the music information retrieval community. Motivated by the rising popularity of virtual and augmented reality experiences as well as potential applications to accessibility, we investigate how well existing music source separation (MSS) models perform on binaural audio. Although these models process two-channel inputs, it is unclear how effectively they retain spatial information. In this work, we evaluate how several popular MSS models preserve spatial information on both standard stereo and novel binaural datasets. Our binaural data is synthesized using stems from MUSDB18-HQ and open-source head-related transfer functions by positioning instrument sources randomly along the horizontal plane. We then assess the spatial quality of the separated stems using signal processing and interaural cue-based metrics. Our results show that stereo MSS models fail to preserve the spatial information critical for maintaining the immersive quality of binaural audio, and that the degradation depends on model architecture as well as the target instrument. Finally, we highlight valuable opportunities for future work at the intersection of MSS and immersive audio.

音乐分离双耳音频空间信息沉浸式音频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。