提出新方法提升虚拟现实3D模型注意力标注精度。
Robust Mesh Saliency Ground Truth Acquisition in VR via View Cone Sampling and Manifold Diffusion
- 用视锥采样模拟人眼聚焦,增强复杂模型采样鲁棒性。
- 融合流形与欧氏约束的扩散算法,避免跨区域误传播。
- 适合关注3D视觉感知、虚拟现实渲染优化的研究者。
随着3D数字内容复杂度指数级增长,理解人类视觉注意力对优化渲染与处理资源至关重要。因此,可靠的3D网格显著性真实标签(GT)对于虚拟现实(VR)中以人为本的视觉建模不可或缺。然而,现有VR眼动追踪框架在采集与生成机制上存在根本瓶颈。依赖零面积单射线采样(SRS)无法捕捉上下文特征,导致严重纹理混叠和显著性信号不连续;而传统欧氏平滑方法会跨物理断开区域传播显著性,造成复杂3D流形上的语义混淆。本文提出一种稳健框架以解决上述问题。首先引入视锥采样(VCS)策略,通过高斯分布射线束模拟人眼中央凹感受野,提升复杂拓扑结构下的采样鲁棒性。进一步提出混合流形-欧氏约束扩散(HCD)算法,融合流形测地线约束与欧氏尺度,确保显著性传播的拓扑一致性。通过主观实验及定性定量分析,验证了本方法在性能上的提升及其对下游任务的益处。该框架有效缓解了‘拓扑短路’与混叠问题,提供了一种符合自然人类感知的高保真3D注意力获取范式,为3D网格显著性研究提供了更准确、更鲁棒的基准。
原文摘要 · Abstract (English)
As the complexity of 3D digital content grows exponentially, understanding human visual attention is critical for optimizing rendering and processing resources. Therefore, reliable 3D mesh saliency ground truth (GT) is essential for human-centric visual modeling in virtual reality (VR). However, existing VR eye-tracking frameworks are fundamentally bottlenecked by their underlying acquisition and generation mechanisms. The reliance on zero-area single ray sampling (SRS) fails to capture contextual features, leading to severe texture aliasing and discontinuous saliency signals. And the conventional application of Euclidean smoothing propagates saliency across disconnected physical gaps, resulting in semantic confusion on complex 3D manifolds. This paper proposes a robust framework to address these limitations. We first introduce a view cone sampling (VCS) strategy, which simulates the human foveal receptive field via Gaussian-distributed ray bundles to improve sampling robustness for complex topologies. Furthermore, a hybrid Manifold-Euclidean constrained diffusion (HCD) algorithm is developed, fusing manifold geodesic constraints with Euclidean scales to ensure topologically-consistent saliency propagation. We demonstrate the improvement in performance over baseline methods and the benefits for downstream tasks through subjective experiments and qualitative and quantitative methods. By mitigating "topological short-circuits" and aliasing, our framework provides a high-fidelity 3D attention acquisition paradigm that aligns with natural human perception, offering a more accurate and robust baseline for 3D mesh saliency research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。