用关系6D位姿图实现物理约束下的机器人操作,零样本泛化强。
RelAfford6D: Relational 6D Affordance Graphs for Constraint-Driven Robotic Manipulation

- 构建关系6D位姿图,将语义拓扑转为精确物理位姿。
- 在仿真与真实场景中实现零样本成功率超基线,支持跨类别泛化。
- 适合需要高鲁棒性与物理一致性的真实机器人任务。
开放世界机器人操作中,如何连接抽象语义与精确物理控制仍是核心挑战。现有数据驱动策略依赖孤立接触点或隐式可操作性嵌入,缺乏复杂铰接物体所需的严格运动学约束。为此,我们提出RelAfford6D——一种无需训练的全新框架,核心是关系6D可操作性图。给定自由形式指令,系统推断出主要交互部件与其物理锚点之间的语义拓扑;通过视觉基础模型将这些拓扑节点提升为精确的$SE(3)$位姿,从而将下游执行建模为运动学约束满足问题。机器人通过追踪严格定义的物理流形(如旋转或平移轨道)生成连续轨迹,并结合闭环跟踪机制动态重规划以应对扰动。该基于物理的方法在仿真与真实环境中均实现优异的零样本成功率、跨类别泛化能力与执行鲁棒性,显著优于现有数据驱动基线。
原文摘要 · Abstract (English)
Bridging abstract semantics and precise physical control remains a fundamental challenge in open-world robotic manipulation. While recent data-driven policies show promise, their reliance on isolated contact points or latent affordance embeddings lacks the rigorous kinematic constraints necessary for complex articulated objects.To overcome the limitation, we introduce RelAfford6D, a novel training-free framework centered on a Relational 6D Affordance Graph. Given a free-form instruction, our system deduces a semantic topology linking a primary interacting part to its physical anchor. By elevating these topological nodes into precise metric $SE(3)$ poses via vision foundation models, we analytically formulate downstream execution as a kinematic constraint satisfaction problem. The robot synthesizes continuous trajectories by tracking strictly defined physical manifolds (e.g., revolute or prismatic orbits). Coupled with a closed-loop tracking mechanism for dynamic replanning against disturbances, our physically grounded approach achieves superior zero-shot success rates, cross-category generalization and execution robustness in both simulation and the real world environments, outperforming existing data-driven baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。