arXiv:2411.18222eess.AScs.SD2024-11中稿 · er for publication…被引 6

用认知模型提升音频质量评估的泛化能力,尤其针对新型编码失真。

Towards Improved Objective Perceptual Audio Quality Assessment -- Part 1: A Novel Data-Driven Cognitive Model

  • 基于主观数据建模听觉认知,自适应加权失真指标。
  • 在未见过的音频数据上,预测准确率优于现有工具。
  • 适合音频编解码器开发与质量评测场景。

高效音频质量评估对简化音频编解码器开发至关重要。客观评估工具通过算法从主观评分(质量判断的金标准)中预测质量等级。现有工具多采用听觉感知模型提取特征,并结合机器学习与主观评分进行训练,但对未知信号和失真类型泛化能力不足,尤其在非波形保真的参数化编码中表现明显。本文(第一部分)提出扩展ITU-R BS.1387-1推荐标准PEAQ的方法,引入一种新型数据驱动的认知模型以增强预测泛化性。该方法利用主观数据建模音频质量感知的认知层面,通过交互代价函数捕捉失真显著性与认知影响的关系,自适应调整不同失真度量的权重。相比其他机器学习方法及成熟工具,该架构在大规模未见主观评分数据库上实现更高预测准确率。所提感知驱动模型比通用机器学习算法更易扩展,支持多维度质量测量的改进而无需完全重训。

原文摘要 · Abstract (English)

Efficient audio quality assessment is vital for streamlining audio codec development. Objective assessment tools have been developed over time to algorithmically predict quality ratings from subjective assessments, the gold standard for quality judgment. Many of these tools use perceptual auditory models to extract audio features that are mapped to a basic audio quality score prediction using machine learning algorithms and subjective scores as training data. However, existing tools struggle with generalization in quality prediction, especially when faced with unknown signal and distortion types. This is particularly evident in the presence of signals coded using non-waveform-preserving parametric techniques. Addressing these challenges, this two-part work proposes extensions to the Perceptual Evaluation of Audio Quality (PEAQ - ITU-R BS.1387-1) recommendation. Part 1 focuses on increasing generalization, while Part 2 targets accurate spatial audio quality measurement in audio coding. To enhance prediction generalization, this paper (Part 1) introduces a novel machine learning approach that uses subjective data to model cognitive aspects of audio quality perception. The proposed method models the perceived severity of audible distortions by adaptively weighting different distortion metrics. The weights are determined using an interaction cost function that captures relationships between distortion salience and cognitive effects. Compared to other machine learning methods and established tools, the proposed architecture achieves higher prediction accuracy on large databases of previously unseen subjective quality scores. The perceptually-motivated model offers a more manageable alternative to general-purpose machine learning algorithms, allowing potential extensions and improvements to multi-dimensional quality measurement without complete retraining.

音频质量认知模型机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。