arXiv:2509.16715eess.AScs.LG2025-09

提出专用于空间音频的质量评估模型,能准确预测主观听感。

QASTAnet: A DNN-based Quality Metric for Spatial Audio

  • 结合低层听觉建模与高层认知判断的神经网络架构
  • 在多种音源类型上与主观评分相关性显著高于现有方法
  • 适合用于空间音频编解码器的开发与对比测试

在空间音频技术发展中,可靠的音频质量评估方法至关重要。目前听感测试仍是标准,但耗时且资源消耗大。已有若干预测主观评分的模型,但在真实信号上泛化能力差。本文提出QASTAnet(Quality Assessment for SpaTial Audio network),一种基于深度神经网络、专用于空间音频(包含ambisonics和binaural)的质量评估模型。由于训练数据稀缺,模型设计旨在小样本下即可训练。为此,我们借鉴专家对低层听觉系统的建模,并用神经网络模拟高层认知判断过程。在多种内容类型(语音、音乐、环境声、自由场、混响声)及编解码伪影场景下,与两个参考指标对比,结果表明QASTAnet克服了现有方法的局限。其预测值与主观评分高度相关,是编解码器研发中比较性能的理想候选指标。

原文摘要 · Abstract (English)

In the development of spatial audio technologies, reliable and shared methods for evaluating audio quality are essential. Listening tests are currently the standard but remain costly in terms of time and resources. Several models predicting subjective scores have been proposed, but they do not generalize well to real-world signals. In this paper, we propose QASTAnet (Quality Assessment for SpaTial Audio network), a new metric based on a deep neural network, specialized on spatial audio (ambisonics and binaural). As training data is scarce, we aim for the model to be trainable with a small amount of data. To do so, we propose to rely on expert modeling of the low-level auditory system and use a neurnal network to model the high-level cognitive function of the quality judgement. We compare its performance to two reference metrics on a wide range of content types (speech, music, ambiance, anechoic, reverberated) and focusing on codec artifacts. Results demonstrate that QASTAnet overcomes the aforementioned limitations of the existing methods. The strong correlation between the proposed metric prediction and subjective scores makes it a good candidate for comparing codecs in their development.

空间音频质量评估深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。