arXiv:2506.19404eess.AScs.SD2025-06综述被引 5

让声音模型更懂人耳的空间感知,提升耳机听感真实度。

Loss functions incorporating auditory spatial perception in deep learning -- a review

  • 用空间感知线索替代传统信号差值,优化声音生成
  • 聚焦定位线索(如双耳时差、声强差),忽略混响等特性
  • 适合做沉浸式音频生成的科研与工程人员参考

双耳还原旨在通过耳机提供具有高度感知真实性的沉浸式空间音频。损失函数在优化和评估生成双耳信号的算法中起核心作用。然而,传统的信号相关差异度量往往无法捕捉空间音频质量所必需的感知特性。本文综述了近年来将空间感知线索纳入双耳再现相关的损失函数。重点考察应用于双耳信号的损失函数,这些信号通常源自麦克风录音或Ambisonics信号,排除基于房间冲击响应的损失。依据空间音频质量清单(SAQI),本文强调与声源定位和房间响应相关的感知维度,排除一般频谱-时序属性。文献调查显示,研究主要集中于定位线索(如双耳时间差、双耳强度差,ITDs、ILDs),而混响及其他房间声学特性在损失函数设计中仍较少被关注。近期工作通过估计房间声学参数并构建捕捉房间特征的嵌入表示,显示出其未来融入神经网络训练的潜力。论文最后指出,未来研究应朝向更贴近听觉体验的感知基础损失函数发展。

原文摘要 · Abstract (English)

Binaural reproduction aims to deliver immersive spatial audio with high perceptual realism over headphones. Loss functions play a central role in optimizing and evaluating algorithms that generate binaural signals. However, traditional signal-related difference measures often fail to capture the perceptual properties that are essential to spatial audio quality. This review paper surveys recent loss functions that incorporate spatial perception cues relevant to binaural reproduction. It focuses on losses applied to binaural signals, which are often derived from microphone recordings or Ambisonics signals, while excluding those based on room impulse responses. Guided by the Spatial Audio Quality Inventory (SAQI), the review emphasizes perceptual dimensions related to source localization and room response, while excluding general spectral-temporal attributes. The literature survey reveals a strong focus on localization cues, such as interaural time and level differences (ITDs, ILDs), while reverberation and other room acoustic attributes remain less explored in loss function design. Recent works that estimate room acoustic parameters and develop embeddings that capture room characteristics indicate their potential for future integration into neural network training. The paper concludes by highlighting future research directions toward more perceptually grounded loss functions that better capture the listener's spatial experience.

空间音频损失函数双耳还原

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。