提出可精准评估立体声音频空间定位一致性的客观指标
BINAQUAL: A Full-Reference Objective Localization Similarity Metric for Binaural Audio
- 基于全景声评估方法改造,适配立体声场景的定位相似性度量
- 在五类测试中均与主观听感高度相关,能区分细微空间差异
- 适合音频压缩、传输等环节的空间保真度验证,节省测试成本
空间音频在虚拟现实、增强现实、游戏和电影中通过三维听觉体验提升沉浸感。由于压缩、编码或传输过程可能改变定位线索,确保立体声音频的空间保真度至关重要。尽管主观听辨测试(如MUSHRA)仍是评估空间定位质量的金标准,但耗时且成本高。本文提出BINAQUAL,一种全参考客观度量方法,用于评估立体声音频记录中的定位相似性。该方法将原本用于全景声格式的AMBIQUAL度量扩展至立体声领域。我们在五个关键研究问题上评估BINAQUAL:声源位置变化、角度插值、环绕扬声器布局、音频退化及内容多样性。结果表明,BINAQUAL能有效区分细微的空间差异,且与主观听感高度相关,具备可靠的定位质量评估能力。该指标为立体声音频处理中的空间准确性提供了稳健基准,推动沉浸式音频应用的客观评估发展。
原文摘要 · Abstract (English)
Spatial audio enhances immersion in applications such as virtual reality, augmented reality, gaming, and cinema by creating a three-dimensional auditory experience. Ensuring the spatial fidelity of binaural audio is crucial, given that processes such as compression, encoding, or transmission can alter localization cues. While subjective listening tests like MUSHRA remain the gold standard for evaluating spatial localization quality, they are costly and time-consuming. This paper introduces BINAQUAL, a full-reference objective metric designed to assess localization similarity in binaural audio recordings. BINAQUAL adapts the AMBIQUAL metric, originally developed for localization quality assessment in ambisonics audio format to the binaural domain. We evaluate BINAQUAL across five key research questions, examining its sensitivity to variations in sound source locations, angle interpolations, surround speaker layouts, audio degradations, and content diversity. Results demonstrate that BINAQUAL effectively differentiates between subtle spatial variations and correlates strongly with subjective listening tests, making it a reliable metric for binaural localization quality assessment. The proposed metric provides a robust benchmark for ensuring spatial accuracy in binaural audio processing, paving the way for improved objective evaluations in immersive audio applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。