用2D提示和几何优化,实现高效3D可操作性分割
Task-Aware 3D Affordance Segmentation via 2D Guidance and Geometric Refinement
- 先用2D语义提示定位任务相关区域,再融合3D几何信息细化结果
- 在SceneFun3D上准确率显著优于基线,推理效率更高
- 适合需要理解环境可操作性的机器人交互场景
从自然语言指令中理解3D场景级可操作性对赋予具身智能体在复杂环境中有意义的交互能力至关重要。然而,该任务因需语义推理与空间定位而极具挑战。现有方法多聚焦于物体级可操作性,或仅将2D预测映射至3D,忽视点云中的丰富几何结构信息,且计算开销大。为此,我们提出任务感知的3D场景级可操作性分割框架TASA,通过粗到细的方式联合利用2D语义线索与3D几何推理。为提升检测效率,TASA设计了任务感知的2D可操作性检测模块,基于语言与视觉输入识别可操作点,指导任务相关视图的选择。为充分挖掘3D几何信息,提出3D可操作性精炼模块,将2D语义先验与局部3D几何融合,生成精确且空间连贯的3D可操作性掩码。在SceneFun3D上的实验表明,TASA在场景级可操作性分割的准确率与效率方面均显著优于基线。
原文摘要 · Abstract (English)
Understanding 3D scene-level affordances from natural language instructions is essential for enabling embodied agents to interact meaningfully in complex environments. However, this task remains challenging due to the need for semantic reasoning and spatial grounding. Existing methods mainly focus on object-level affordances or merely lift 2D predictions to 3D, neglecting rich geometric structure information in point clouds and incurring high computational costs. To address these limitations, we introduce Task-Aware 3D Scene-level Affordance segmentation (TASA), a novel geometry-optimized framework that jointly leverages 2D semantic cues and 3D geometric reasoning in a coarse-to-fine manner. To improve the affordance detection efficiency, TASA features a task-aware 2D affordance detection module to identify manipulable points from language and visual inputs, guiding the selection of task-relevant views. To fully exploit 3D geometric information, a 3D affordance refinement module is proposed to integrate 2D semantic priors with local 3D geometry, resulting in accurate and spatially coherent 3D affordance masks. Experiments on SceneFun3D demonstrate that TASA significantly outperforms the baselines in both accuracy and efficiency in scene-level affordance segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。