arXiv:2607.07015cs.SD2026-07

用生成式空间音频帮视障学习者在360度环境中定位和构建空间认知。

EscFOA: Enhancing Spatial Learning for Visually Impaired Learners via Generative Spatial Audio in 360-Degree Educational Environments

论文配图:EscFOA: Enhancing Spatial Learning for Visually Impaired Learners via Generative Spatial Audio in 360-Degree Educational Environments
图 1 · 摘自论文原文
  • 结合3D高斯点云与条件扩散模型重建场景结构,生成与环境一致的空间音频。
  • 在盲测实验中显著优于传统单声道和立体声,提升空间定向能力与学习效率。
  • 适合需要空间认知支持的教育类虚拟现实应用,尤其对视障学习者友好。

沉浸式360度教育环境常缺乏可访问的空间结构,限制了视障学习者的方位感知、探索能力和心理表征构建。本文提出EscFOA,一种几何感知的空间音频生成框架,作为支持空间认知的“听觉支架”。通过融合3D高斯点云(3DGS)与条件扩散模型,EscFOA从360度视频中重建场景几何,并合成与环境结构一致的高保真空间音频。该方法明确针对独立空间定向与降低认知负荷等学习成果,在模拟视障学习者的盲测实验中显著优于传统的单声道与立体声音频,验证了几何一致性生成音频能有效促进复杂空间学习材料的包容性获取。

原文摘要 · Abstract (English)

Immersive 360-degree educational environments often lack accessible spatial structure, limiting visually impaired learners' ability to orient, explore, and construct mental representations. This paper proposes EscFOA, a geometry-aware spatial audio generation framework designed as an \emph{acoustic scaffolding} to support spatial cognition. By integrating 3D Gaussian Splatting (3DGS) with conditional diffusion models, EscFOA reconstructs scene geometry from 360-degree videos to synthesize high-fidelity spatial audio consistent with the environmental structure. Explicitly targeting learning outcomes like independent spatial orientation and reduced cognitive load, EscFOA significantly outperforms conventional monaural and stereo audio in supporting spatial learning behaviors among blindfolded sighted participants (simulating visually impaired learners). These findings demonstrate that geometry-consistent generative audio can effectively enable inclusive access to complex spatial learning materials.

空间音频视障教育360度环境生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。