arXiv:2602.06419cs.CV2026-02

用几何引导语义先验,让模型更像人一样看3D物体。

Learning Human Visual Attention on 3D Surfaces through Geometry-Queried Semantic Priors

  • 双流结构让几何与语义信息不对称融合,模拟人类视觉机制。
  • 在SAL3D等3个数据集上显著提升预测准确率,尤其在语义重要但几何平庸区域表现更好。
  • 适合关注3D视觉建模、认知神经科学或人机交互的研究者。

三维物体上的视觉注意源于自下而上几何处理与自上而下语义识别的相互作用。现有3D显著性方法依赖手工设计的几何特征或缺乏语义感知的学习方法,无法解释为何人类会注视语义重要但几何不显著的区域。本文提出SemGeo-AttentionNet,一种双流架构,通过非对称跨模态融合显式建模这一二元性:利用基于扩散的语义先验(来自几何条件的多视图渲染)和点云变换器进行几何处理。交叉注意力使几何特征主动查询语义内容,实现自下而上显著性引导自上而下检索。进一步通过强化学习扩展至时序扫描路径生成,提出首个尊重3D网格拓扑并包含回抑制动态的建模方法。在SAL3D、NUS3D和3DVA数据集上的评估表明性能显著提升,验证了认知启发架构在建模人类3D表面视觉注意方面的有效性。

原文摘要 · Abstract (English)

Human visual attention on three-dimensional objects emerges from the interplay between bottom-up geometric processing and top-down semantic recognition. Existing 3D saliency methods rely on hand-crafted geometric features or learning-based approaches that lack semantic awareness, failing to explain why humans fixate on semantically meaningful but geometrically unremarkable regions. We introduce SemGeo-AttentionNet, a dual-stream architecture that explicitly formalizes this dichotomy through asymmetric cross-modal fusion, leveraging diffusion-based semantic priors from geometry-conditioned multi-view rendering and point cloud transformers for geometric processing. Cross-attention ensures geometric features query semantic content, enabling bottom-up distinctiveness to guide top-down retrieval. We extend our framework to temporal scanpath generation through reinforcement learning, introducing the first formulation respecting 3D mesh topology with inhibition-of-return dynamics. Evaluation on SAL3D, NUS3D and 3DVA datasets demonstrates substantial improvements, validating how cognitively motivated architectures effectively model human visual attention on three-dimensional surfaces.

3D视觉注意力模型语义先验认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。