解决古典音乐小合奏源分离难题,提升助听设备听感体验
Source Separation of Small Classical Ensembles: Challenges and Opportunities
- 用卷积TasNet构建因果与非因果模型,对比处理效果
- 真实数据上因果/非因果性能相近(0.3 dB vs 0.4 dB SDR)
- 合成数据与真实录音存在差距,需改善生成真实感
使用非因果深度学习进行西方流行音乐的音乐源分离(MSS)已十分有效,但古典音乐的MSS仍是未解难题。原因包括音乐本身变化更大、带真实标签的录音稀缺,以及乐器间区分度更低。为推动研究,Cadenza项目创建了新的合成木管乐合奏数据库,以弥补EnsembleSet中的乐器不平衡问题。采用一组ConvTasNet模型分别提取弦乐或木管乐器,因其支持因果与非因果两种处理方式。非因果方法在录制音乐中表现优异,但对助听器等实时场景,需依赖因果处理。在真实乐器录音数据集Bach10和URMP上评估性能,发现因果与非因果系统表现接近(平均SDR分别为0.3 dB和0.4 dB)。合成验证集的性能更好(6.2 dB因果,6.9 dB非因果),表明合成数据与真实录音之间存在显著差异。未来工作需收集更多真实录音或提升合成数据的真实感与多样性。
原文摘要 · Abstract (English)
Musical (MSS) source separation of western popular music using non-causal deep learning can be very effective. In contrast, MSS for classical music is an unsolved problem. Classical ensembles are harder to separate than popular music because of issues such as the inherent greater variation in the music; the sparsity of recordings with ground truth for supervised training; and greater ambiguity between instruments. The Cadenza project has been exploring MSS for classical music. This is being done so music can be remixed to improve listening experiences for people with hearing loss. To enable the work, a new database of synthesized woodwind ensembles was created to overcome instrumental imbalances in the EnsembleSet. For the MSS, a set of ConvTasNet models was used with each model being trained to extract a string or woodwind instrument. ConvTasNet was chosen because it enabled both causal and non-causal approaches to be tested. Non-causal approaches have dominated MSS work and are useful for recorded music, but for live music or processing on hearing aids, causal signal processing is needed. The MSS performance was evaluated on the two small datasets (Bach10 and URMP) of real instrument recordings where the ground-truth is available. The performances of the causal and non-causal systems were similar. Comparing the average Signal-to-Distortion (SDR) of the synthesized validation set (6.2 dB causal; 6.9 non-causal), to the real recorded evaluation set (0.3 dB causal, 0.4 dB non-causal), shows that mismatch between synthesized and recorded data is a problem. Future work needs to either gather more real recordings that can be used for training, or to improve the realism and diversity of the synthesized recordings to reduce the mismatch...
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。