arXiv:2409.14069eess.AScs.SD2024-09中稿 · ICASSP 2025

将音频评估转化为多模态文本预测,模拟人类专注听觉注意力。

Semi-intrusive audio evaluation: Casting non-intrusive assessment as a multi-modal text prediction task

  • 用音频+文本输入训练模型,实现类人注意力聚焦。
  • 在MOS评分上相关性提升0.06至0.20,优于现有方法。
  • 新提出的信噪比估计算法可精准定位特定声源,适合音效分析场景。

人类感知具有在混合信号中聚焦特定事件的独特能力,而现有非侵入式评估方法难以应对这一挑战。本文提出半侵入式评估,通过将音频评估建模为音频-文本输入的文本预测任务,模拟人类注意力机制。我们对多模态PENGI模型进行指令微调,用于MOS和信噪比(SNR)估计。在MOS任务中,该方法相比重新训练的MOSRA模型和预训练的PAM模型,分别获得0.06和0.20的皮尔逊相关系数绝对提升。我们还提出一种新型SNR估计算法,能聚焦于混合音频中的特定声源,显著优于随机基线和固定提示对照组。结果表明,半侵入式评估可有效捕捉人类选择性听觉能力。样本可在 https://jozefcoldenhoff.github.io/semi-intrusive-assessment 获取。

原文摘要 · Abstract (English)

Human perception has the unique ability to focus on specific events in a mixture of signals--a challenging task for existing non-intrusive assessment methods. In this work, we introduce semi-intrusive assessment that emulates human attention by framing audio assessment as a text-prediction task with audio-text inputs. To this end, we extend the multi-modal PENGI model through instruction fine-tuning for MOS and SNR estimation. For MOS, our approach achieves absolute Pearson correlation gains of 0.06 and 0.20 over the re-trained MOSRA model and the pre-trained PAM model, respectively. We further propose a novel SNR estimator that can focus on a specific audio source in a mixture, outperforming a random baseline and the fixed-prompt counterpart. Our findings suggest that semi-intrusive assessment can effectively capture human-like selective listening capabilities. Samples are available at https://jozefcoldenhoff.github.io/semi-intrusive-assessment.

音频评估多模态注意力机制语音质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。