arXiv:2609.06488cs.SDeess.AS2026-09

利用发音音素对齐提升人声主唱分离效果

Lead Vocal Separation from Vocal Ensemble Mixtures Using Phoneme Alignment

论文配图:Lead Vocal Separation from Vocal Ensemble Mixtures Using Phoneme Alignment
图 1 · 摘自论文原文
  • 通过音素对齐信息增强模型中间表示,引导主唱分离
  • 相比仅用演唱/静音状态,音素条件使性能提升更显著
  • 适合需要精准主唱提取的音乐处理与伴奏生成场景

当代无伴奏合唱常呈现主唱与伴唱结合的结构,其中主唱(Vo)承担主要旋律,其余声部提供和声伴奏。将主唱从合唱中分离(即Vo分离)可支持歌词识别、零一伴奏生成等下游应用。然而,由于目标与干扰源均为人声且声学特征相似、时间上高度重叠,该任务面临线索稀缺的挑战。本文提出一种基于音素对齐信息的Vo分离模型,以带频段分割的RoPE Transformer(BS-RoFormer)为基础,通过特征逐维线性调制(FiLM)将帧级音素标签引入中间表征。实验表明,音素对齐条件能显著优于纯音频基线,并在主唱与其他声部共享相同音素较少时带来更大增益。

原文摘要 · Abstract (English)

Contemporary a cappella singing often has a lead-and-accompaniment texture, where the lead vocal (Vo) part carries the main melody and the remaining vocal parts provide accompaniment. Owing to their distinct roles, separating the Vo part from the remaining vocal parts, referred to as Vo separation, enables downstream applications such as lyric recognition and minus-one accompaniment generation for vocal ensemble music. Despite these potential applications, acoustic cues for this task are limited because the target and interfering sources are all singing voices with similar acoustic characteristics and often overlap in time, making Vo separation challenging. In this paper, we propose a Vo separation model that uses phoneme alignment of the Vo part as auxiliary information. The proposed model is based on band-split RoPE Transformer (BS-RoFormer), a state-of-the-art music source separation model, and introduces frame-level phoneme labels into its intermediate representations using feature-wise linear modulation (FiLM). Experimental results show that phoneme-alignment conditioning improves Vo separation performance over an audio-only baseline and yields larger average gains than conditioning only on Vo singing/silence activity. Further analysis suggests that the advantage of phoneme-label information is larger when fewer remaining vocal parts share the same phoneme as Vo.

人声分离音素对齐音乐处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。