arXiv:2509.00400eess.AScs.SD2025-09被引 6

用深度学习实现个性化立体声播放,提升沉浸感与听觉真实感

Deep Learning for Personalized Binaural Audio Reproduction

  • 分两类:显式预测个人化头相关传输函数,或端到端直接生成立体声信号
  • 结合形态特征、环境线索等信息,提升空间定位精度和外化效果
  • 适合做虚拟现实、助听器等需要精准空间音频的场景

个性化立体声音频还原是实现真实空间定位、声音外化和沉浸式聆听的基础,直接影响用户体验和听觉负担。本文综述了深度学习在该任务中的最新进展,按生成机制分为两类:显式个性化滤波与端到端渲染。显式方法从稀疏测量、形态特征或环境线索中预测个性化头相关传输函数(HRTFs),并用于传统渲染流程;端到端方法则在视觉、文本或参数引导下,直接将声源信号映射为立体声信号,并在模型内部学习个性化。本文还总结了该领域的主流数据集与评估指标,以支持公平可复现的对比。最后讨论了这些技术推动的关键应用、当前技术局限及未来研究方向。

原文摘要 · Abstract (English)

Personalized binaural audio reproduction is the basis of realistic spatial localization, sound externalization, and immersive listening, directly shaping user experience and listening effort. This survey reviews recent advances in deep learning for this task and organizes them by generation mechanism into two paradigms: explicit personalized filtering and end-to-end rendering. Explicit methods predict personalized head-related transfer functions (HRTFs) from sparse measurements, morphological features, or environmental cues, and then use them in the conventional rendering pipeline. End-to-end methods map source signals directly to binaural signals, aided by other inputs such as visual, textual, or parametric guidance, and they learn personalization within the model. We also summarize the field's main datasets and evaluation metrics to support fair and repeatable comparison. Finally, we conclude with a discussion of key applications enabled by these technologies, current technical limitations, and potential research directions for deep learning-based spatial audio systems.

空间音频深度学习个性化虚拟现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。