用动态语义流让机器人更懂物体,实现精准且通用的抓取操作。
G3Flow: Generative 3D Semantic Flow for Pose-aware and Generalizable Object Manipulation
- 构建实时3D语义流,融合视觉大模型与生成模型理解物体
- 在遮挡下仍保持高精度,终端约束任务成功率最高达68.3%
- 适合需要复杂场景理解的机器人操控研究者
近期基于扩散模型的3D机器人操作模仿学习取得进展,但要达到人类级灵巧性,需融合几何精度与语义理解。我们提出G3Flow,一种通过基础模型构建实时语义流的新框架,该流为以物体为中心的动态3D语义表示。方法结合3D生成模型创建数字孪生、视觉基础模型提取语义特征,并利用鲁棒位姿跟踪实现连续语义流更新。该集成在遮挡条件下仍保持完整语义理解,无需人工标注。将语义流引入扩散策略后,在五项仿真任务中显著提升末端约束操作与跨物体泛化能力。实验表明,G3Flow在终端约束任务上平均成功率最高达68.3%,跨物体泛化任务达50.1%。结果证明了该框架在增强机器人操作策略实时动态语义理解方面的有效性。
原文摘要 · Abstract (English)
Recent advances in imitation learning for 3D robotic manipulation have shown promising results with diffusion-based policies. However, achieving human-level dexterity requires seamless integration of geometric precision and semantic understanding. We present G3Flow, a novel framework that constructs real-time semantic flow, a dynamic, object-centric 3D semantic representation by leveraging foundation models. Our approach uniquely combines 3D generative models for digital twin creation, vision foundation models for semantic feature extraction, and robust pose tracking for continuous semantic flow updates. This integration enables complete semantic understanding even under occlusions while eliminating manual annotation requirements. By incorporating semantic flow into diffusion policies, we demonstrate significant improvements in both terminal-constrained manipulation and cross-object generalization. Extensive experiments across five simulation tasks show that G3Flow consistently outperforms existing approaches, achieving up to 68.3% and 50.1% average success rates on terminal-constrained manipulation and cross-object generalization tasks respectively. Our results demonstrate the effectiveness of G3Flow in enhancing real-time dynamic semantic feature understanding for robotic manipulation policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。