arXiv:2509.12712cs.SDcs.IR2025-09

轻量级模型实现任意乐器的动态分离与联合音符转录

A Lightweight Two-Branch Architecture for Multi-Instrument Transcription via Note-Level Contrastive Clustering

  • 双分支结构:通用声部骨干+专用音色编码器
  • 音符级对比聚类,支持任意数量乐器分离
  • 小模型快推理,适合低资源设备部署

现有多音色转录模型在预训练乐器外泛化能力差、源数限制僵化、计算开销大,难以在低资源设备上部署。本文提出轻量级双分支架构,扩展无音色依赖的转录主干,加入专用音色编码器,并在音符层级进行对比聚类,实现给定乐器类别数下的联合转录与动态分离。通过谱归一化、空洞卷积及对比聚类等优化,显著提升效率与鲁棒性。尽管模型规模小、推理快,其转录准确率与分离质量仍可媲美更重的基线模型,展现出良好泛化能力,适用于实际场景中资源受限的部署环境。

原文摘要 · Abstract (English)

Existing multi-timbre transcription models struggle with generalization beyond pre-trained instruments, rigid source-count constraints, and high computational demands that hinder deployment on low-resource devices. We address these limitations with a lightweight model that extends a timbre-agnostic transcription backbone with a dedicated timbre encoder and performs deep clustering at the note level, enabling joint transcription and dynamic separation of arbitrary instruments given a specified number of instrument classes. Practical optimizations including spectral normalization, dilated convolutions, and contrastive clustering further improve efficiency and robustness. Despite its small size and fast inference, the model achieves competitive performance with heavier baselines in terms of transcription accuracy and separation quality, and shows promising generalization ability, making it highly suitable for real-world deployment in practical and resource-constrained settings.

音乐转录轻量模型聚类多乐器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。