arXiv:2605.30469cs.SDcs.CV2026-05

提出3DAE框架,用空间误差图诊断音频新视角合成的失败原因。

3DAE: Binaural Quality Assessment for Audio Novel View Synthesis with Spatial Maps and Benchmark

论文配图:3DAE: Binaural Quality Assessment for Audio Novel View Synthesis with Spatial Maps and Benchmark
图 1 · 摘自论文原文
  • 通过时频域误差图分析幅度、声像差、相位差等六类空间音频问题
  • 在Replay-NVAS和SoundSpaces数据集上分别发现时序错位和声像差不匹配为主因
  • 模型无关的诊断工具,适合优化音频新视角合成模型的开发

3D音频与新视角声学合成模型通常依赖全局指标评估,但这类指标常掩盖具体失败位置与原因。本文提出一种全参考诊断框架,利用时频域音频误差图分析幅度、左右耳强度差(ILD)、左右耳相位差(IPD)、时间对齐、响度及高频失真等六类问题,构建三维音频误差图(3DAE Map)用于可视化检视。将该诊断方法整合为一个模型无关的基准测试系统——空间音频误差基准(3DAE Bench),可接收任意真实值与预测的双耳音频对,报告新视角音频合成模型的性能。在ViGAS模型于Replay-NVAS与SoundSpaces数据集上的实验表明:Replay-NVAS中主要失败模式为时间对齐错误,SoundSpaces中则以ILD不匹配为主。整体框架提供可解释的失败模式总结与直观的视觉化地图,助力音频新视角合成模型的优化开发。

原文摘要 · Abstract (English)

3D audio and novel-view acoustic synthesis models are usually evaluated with global metrics.However, global metrics often hide where and why binaural prediction fails. We propose a full-reference diagnostic framework that uses time-frequency audio error maps for magnitude, ILD, IPD, temporal alignment, loudness, and high-frequency failures, forming a 3D Audio Error Map (3DAE Map) for visual inspection. We frame these diagnostics into a model-agnostic benchmark, Spatial Audio Error Bench (3DAE Bench), which takes arbitrary ground-truth and predicted binaural pairs and reports the prediction quality of audio novel-view synthesis models. Experiments on ViGAS outputs over Replay-NVAS and SoundSpaces show different dominant failure modes: temporal misalignment on Replay-NVAS and ILD mismatch on SoundSpaces. Overall, the framework provides interpretable failure-mode summaries and intuitive visual maps for audio Novel-view-synthesis model development optimization.

音频合成空间音频误差分析可视化诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。