让智能体通过交互学习3D可动物体的结构与功能
IAAO: Interactive Affordance Learning for Articulated Objects in 3D Environments
- 用3D高斯点云构建物体层次化特征和标签场
- 分阶段识别静态/可动部分,估计运动参数与交互属性
- 适合需要真实交互的机器人操控与场景理解任务
本文提出IAAO框架,通过交互使智能体在3D环境中理解可动物体。不同于依赖特定任务网络或移动部件假设的方法,IAAO利用大模型分三阶段估计交互属性与部件运动。首先,基于多视角图像,通过3D高斯点云(3DGS)蒸馏掩码特征与视图一致标签,构建每个物体状态的层次化特征与标签场。其次,对3D高斯原语进行物体级与部件级查询,识别静止与可动元素,估计全局变换与局部运动参数及交互属性。最后,基于估计的变换融合并优化不同状态下的场景,实现鲁棒的基于属性的交互与操作。实验验证了方法的有效性。
原文摘要 · Abstract (English)
This work presents IAAO, a novel framework that builds an explicit 3D model for intelligent agents to gain understanding of articulated objects in their environment through interaction. Unlike prior methods that rely on task-specific networks and assumptions about movable parts, our IAAO leverages large foundation models to estimate interactive affordances and part articulations in three stages. We first build hierarchical features and label fields for each object state using 3D Gaussian Splatting (3DGS) by distilling mask features and view-consistent labels from multi-view images. We then perform object- and part-level queries on the 3D Gaussian primitives to identify static and articulated elements, estimating global transformations and local articulation parameters along with affordances. Finally, scenes from different states are merged and refined based on the estimated transformations, enabling robust affordance-based interaction and manipulation of objects. Experimental results demonstrate the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。