arXiv:2606.01460cs.SDeess.AS2026-06

用轻量槽注意力模型实现多乐器多音高检测,区分音高来源乐器。

A Lightweight Slot-Attention Framework for Multi-Instrument Multi-Pitch Estimation

论文配图:A Lightweight Slot-Attention Framework for Multi-Instrument Multi-Pitch Estimation
图 1 · 摘自论文原文
  • 通过无序槽机制与匈牙利匹配,让模型自适应识别乐器音高组合。
  • 在URMP数据集上,匈牙利匹配显著提升乐器族分解准确率。
  • 适合关注源感知音高检测与轻量模型设计的研究者。

多音高估计(MPE)通常仅预测混合音频中哪些音高活跃,但无法判断由何种乐器产生。本文提出一种轻量级槽注意力框架用于多乐器多音高估计(MI-MPE),将混合音频的CQT特征映射为一组无序的源级音高图。模型采用排列不变的匈牙利匹配机制,避免输出语义固定,并将槽数设为活跃声源数量的上限。进一步研究了两项模块化扩展:自监督音色编码器提供槽级音色嵌入的训练目标,以及音符密度正则化分支,约束混合与槽级预测的音高密度。实验表明,匈牙利匹配显著提升了URMP数据集上的乐器族分解性能。但在音轨级预测上仍具挑战:音色与音符密度监督虽改善部分配置,却未稳定解决声源分配问题。结果表明,基于槽的架构是源感知多音高估计的有前景方向,但需更精细地关联辅助音乐线索与槽身份。

原文摘要 · Abstract (English)

Multi-pitch estimation (MPE) typically predicts which pitches are active in a mixture, but not which instrument or source produced them. This paper investigates a lightweight slot-attention framework for multi-instrument MPE (MI-MPE), where a mixture CQT is mapped to an unordered set of source-like pitch maps. The model uses permutation-invariant Hungarian matching to avoid fixed output semantics and treats the number of slots as an upper bound on the number of active sources. We further study two modular extensions: a self-supervised timbre encoder that provides training-time targets for slot-level timbre embeddings, and a polyphony branch that regularizes the pitch density of mixture- and slot-level predictions. Experiments show that Hungarian matching substantially improves instrument family decomposition on URMP. Stem-level prediction remains more challenging: timbre and polyphony supervision improve selected configurations, but do not consistently resolve source assignment. The results suggest that slot-based architectures are a promising direction for source-aware MPE, while highlighting the need to couple auxiliary musical cues to slot identity more carefully.

多音高估计槽注意力乐器分离轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。