arXiv:2503.07828cs.CV2025-03被引 1

用3D神经辐射场建模视觉注意力,实时生成注视热点图。

Neural Radiance and Gaze Fields for Visual Attention Modeling in 3D Environments

  • 在NeRF基础上加注视概率网络,根据场景几何和观察者位置预测注视分布。
  • 实现每秒24帧的交互式注视场渲染,支持视角与摄像机分离。
  • 适用于人机交互、VR/AR中的注意力分析,尤其适合带姿态追踪的数据。

我们提出神经辐射与注视场(NeRGs),一种在复杂环境中表示视觉注意力的新方法。类似于神经辐射场(NeRF)进行新视角合成,NeRGs能从任意视角重建注视模式,将视觉注意力隐式映射到3D表面。通过在标准NeRF上增加一个额外网络,该网络基于场景几何和观察者位置建模局部自我中心注视概率密度。一个NeRG的输出是场景渲染图像与像素级显著性图,表示特定观察者在可见表面上固定注视的条件概率。与以往方法不同,本系统轻量且支持交互式帧率下的注视场可视化。此外,NeRGs可解耦观察者视角与渲染相机,并正确处理因中间几何遮挡导致的注视遮蔽问题。我们使用骨骼追踪获取的头部姿态作为注视代理,通过提出的注视探测器将噪声光线聚合为稳健的概率密度目标以供监督。

原文摘要 · Abstract (English)

We introduce Neural Radiance and Gaze Fields (NeRGs), a novel approach for representing visual attention in complex environments. Much like how Neural Radiance Fields (NeRFs) perform novel view synthesis, NeRGs reconstruct gaze patterns from arbitrary viewpoints, implicitly mapping visual attention to 3D surfaces. We achieve this by augmenting a standard NeRF with an additional network that models local egocentric gaze probability density, conditioned on scene geometry and observer position. The output of a NeRG is a rendered view of the scene alongside a pixel-wise salience map representing the conditional probability that a given observer fixates on visible surfaces. Unlike prior methods, our system is lightweight and enables visualization of gaze fields at interactive framerates. Moreover, NeRGs allow the observer perspective to be decoupled from the rendering camera and correctly account for gaze occlusion due to intervening geometry. We demonstrate the effectiveness of NeRGs using head pose from skeleton tracking as a proxy for gaze, employing our proposed gaze probes to aggregate noisy rays into robust probability density targets for supervision.

视觉注意力3D建模神经辐射场注视预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。