arXiv:2409.05095cs.SDcs.LG2024-09被引 8

用机器学习优化听力障碍者听音乐体验,通过音轨分离重混提升可听性。

The first Cadenza challenges: using machine learning competitions to improve music for listeners with a hearing loss

  • 设计开放竞赛,利用机器学习分离并个性化重混流行摇滚音乐。
  • ICASSP24挑战中9个参赛方案优于基线,最高分采用多模型集成策略。
  • 提供开源数据与代码,推动听障音乐感知研究发展。

听力损失者听音乐存在困难,助听器并非万能解决方案。本文首次将开放竞赛方法应用于机器学习改善听障人群的音乐音频质量。第一项挑战(CAD1)有9名参赛者,第二项为2024年ICASSP大会主竞赛(ICASSP24),吸引17人参与。任务聚焦于流行/摇滚音乐的音轨分离与重混,实现乐器音量个性化调整,并结合放大以补偿听力阈值升高。参赛者基于两大先进分离算法:Hybrid Demucs 和 Open-Unmix。评估使用客观指标 HAAQI(Hearing-Aid Audio Quality Index)。CAD1中无人超越最佳基线,因改进空间有限。因此,ICASSP24采用扬声器播放并设定预重混增益,使场景更贴近助听器使用。9名参赛者表现优于最优基线。多数使用改进版 Hybrid Demucs 与 NAL-R 放大方案。最高分系统采用多个分离算法输出的集成策略。相关软件与数据已公开,形成未来研究基准。

原文摘要 · Abstract (English)

It is well established that listening to music is an issue for those with hearing loss, and hearing aids are not a universal solution. How can machine learning be used to address this? This paper details the first application of the open challenge methodology to use machine learning to improve audio quality of music for those with hearing loss. The first challenge was a stand-alone competition (CAD1) and had 9 entrants. The second was an 2024 ICASSP grand challenge (ICASSP24) and attracted 17 entrants. The challenge tasks concerned demixing and remixing pop/rock music to allow a personalised rebalancing of the instruments in the mix, along with amplification to correct for raised hearing thresholds. The software baselines provided for entrants to build upon used two state-of-the-art demix algorithms: Hybrid Demucs and Open-Unmix. Evaluation of systems was done using the objective metric HAAQI, the Hearing-Aid Audio Quality Index. No entrants improved on the best baseline in CAD1 because there was insufficient room for improvement. Consequently, for ICASSP24 the scenario was made more difficult by using loudspeaker reproduction and specified gains to be applied before remixing. This also made the scenario more useful for listening through hearing aids. 9 entrants scored better than the the best ICASSP24 baseline. Most entrants used a refined version of Hybrid Demucs and NAL-R amplification. The highest scoring system combined the outputs of several demixing algorithms in an ensemble approach. These challenges are now open benchmarks for future research with the software and data being freely available.

听障音乐音轨分离机器学习开放竞赛

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。