提出全景物体使用性识别新框架,解决360°环境下的语义漂移问题。
PanoAffordanceNet: Towards Holistic Affordance Grounding in 360° Indoor Environments
- 设计畸变感知调制器与全景稠密化头,应对投影畸变和拓扑断裂
- 在360-AGD数据集上达到新高,显著优于现有方法
- 适合做全景智能体场景理解的研究者与开发者
全局感知对360°空间中的具身智能体至关重要,但当前的使用性定位仍以物体为中心且局限于视角视图。为此,我们提出全新任务:360°室内环境中的整体使用性定位。该任务面临严重几何畸变(来自等距柱状投影)、语义分散及跨尺度对齐困难等挑战。我们提出PanoAffordanceNet,一个端到端框架,包含畸变感知谱调制器(DASM)实现纬度依赖校准,以及全向球面稠密化头(OSDH)从稀疏激活中恢复拓扑连续性。通过融合像素级、分布级与区域-文本对比约束,框架在弱监督下有效抑制了语义漂移。此外,我们构建了首个高质量全景使用性定位数据集360-AGD。大量实验表明,PanoAffordanceNet显著优于现有方法,为具身智能中的场景级感知建立了坚实基线。源代码与基准数据集将公开于https://github.com/GL-ZHU925/PanoAffordanceNet。
原文摘要 · Abstract (English)
Global perception is essential for embodied agents in 360° spaces, yet current affordance grounding remains largely object-centric and restricted to perspective views. To bridge this gap, we introduce a novel task: Holistic Affordance Grounding in 360° Indoor Environments. This task faces unique challenges, including severe geometric distortions from Equirectangular Projection (ERP), semantic dispersion, and cross-scale alignment difficulties. We propose PanoAffordanceNet, an end-to-end framework featuring a Distortion-Aware Spectral Modulator (DASM) for latitude-dependent calibration and an Omni-Spherical Densification Head (OSDH) to restore topological continuity from sparse activations. By integrating multi-level constraints comprising pixel-wise, distributional, and region-text contrastive objectives, our framework effectively suppresses semantic drift under low supervision. Furthermore, we construct 360-AGD, the first high-quality panoramic affordance grounding dataset. Extensive experiments demonstrate that PanoAffordanceNet significantly outperforms existing methods, establishing a solid baseline for scene-level perception in embodied intelligence. The source code and benchmark dataset will be made publicly available at https://github.com/GL-ZHU925/PanoAffordanceNet.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。