用全景图像提升智能体对环境功能的全局理解能力
Panoramic Affordance Prediction
- 模仿人类视觉聚焦机制,分步定位目标并校正畸变
- 在12000+张全景图上实现精准实例级功能标注
- 适合需要全局空间感知的机器人、自动驾驶场景
affordance预测是具身智能中连接感知与行动的关键桥梁。然而,现有研究局限于针孔相机模型,存在视场窄、观测碎片化等问题,常遗漏关键的整体环境信息。本文首次探索全景功能预测,利用360度图像捕捉全局空间关系与整体场景理解。为推动该任务,我们构建了PAP-12K大规模基准数据集,包含超过1000张超高清(12k,11904×5952)全景图像,以及超过12000个精心标注的问答对和功能掩码。此外,提出PAP框架——一种无需训练、类人视觉聚焦的粗到精流程,通过网格提示递归视觉路由逐步定位目标,采用自适应凝视机制校正局部几何畸变,并使用级联定位流程提取精确的实例级掩码。在PAP-12K上的实验表明,针对标准透视图像设计的现有方法因全景视觉的独特挑战而性能严重退化甚至失效;而PAP框架有效克服这些障碍,显著优于最先进基线,凸显全景感知在鲁棒具身智能中的巨大潜力。
原文摘要 · Abstract (English)
Affordance prediction serves as a critical bridge between perception and action in embodied AI. However, existing research is confined to pinhole camera models, which suffer from narrow Fields of View (FoV) and fragmented observations, often missing critical holistic environmental context. In this paper, we present the first exploration into Panoramic Affordance Prediction, utilizing 360-degree imagery to capture global spatial relationships and holistic scene understanding. To facilitate this novel task, we first introduce PAP-12K, a large-scale benchmark dataset containing over 1,000 ultra-high-resolution (12k, 11904 x 5952) panoramic images with over 12k carefully annotated QA pairs and affordance masks. Furthermore, we propose PAP, a training-free, coarse-to-fine pipeline inspired by the human foveal visual system to tackle the ultra-high resolution and severe distortion inherent in panoramic images. PAP employs recursive visual routing via grid prompting to progressively locate targets, applies an adaptive gaze mechanism to rectify local geometric distortions, and utilizes a cascaded grounding pipeline to extract precise instance-level masks. Experimental results on PAP-12K reveal that existing affordance prediction methods designed for standard perspective images suffer severe performance degradation and fail due to the unique challenges of panoramic vision. In contrast, PAP framework effectively overcomes these obstacles, significantly outperforming state-of-the-art baselines and highlighting the immense potential of panoramic perception for robust embodied intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。