分阶段解耦抓取与操作,提升机器人对异类物体的通用操控能力。
HeteroGenManip: Generalizable Manipulation For Heterogeneous Object Interactions

- 先定位接触点,再分类型调用专用模型执行操作
- 仿真中平均性能提升31%,真实任务提升36.7%
- 适合需要跨类别交互的复杂机械操作场景
涉及异类物体交互的通用化操作是机器人领域关键且具挑战性的能力。为可靠完成此类任务,机器人需解决两个核心问题:‘何处操作’(接触点定位)与‘如何操作’(后续交互轨迹规划)。现有基于基础模型的方法多采用端到端学习,模糊了两阶段界限,导致长时序任务中误差累积加剧;且通常依赖单一统一模型,难以捕捉异类物体所需的多样化、类别特异性特征。为此,我们提出HeteroGenManip,一种任务条件化的两阶段框架,旨在解耦初始抓取与复杂交互执行。首先,通过结构先验引导的基座对应抓取模块,对齐初始接触状态,显著降低抓取位姿不确定性。随后,多基础模型扩散策略(MFMDP)将物体路由至类别专有的基础模型,利用双流交叉注意力机制融合细粒度几何信息与高度可变的部件特征。实验表明,该框架在仿真中实现跨类别形状与位姿的稳健泛化,在广泛类型设置下平均性能提升31%;在四种不同交互类型的四次真实任务中,性能提升达36.7%。
原文摘要 · Abstract (English)
Generalizable manipulation involving cross-type object interactions is a critical yet challenging capability in robotics. To reliably accomplish such tasks, robots must address two fundamental challenges: "where to manipulate" (contact point localization) and "how to manipulate" (subsequent interaction trajectory planning). Existing foundation-model-based approaches often adopt end-to-end learning that obscures the distinction between these stages, exacerbating error accumulation in long-horizon tasks. Furthermore, they typically rely on a single uniform model, which fails to capture the diverse, category-specific features required for heterogeneous objects. To overcome these limitations, we propose HeteroGenManip, a task-conditioned, two-stage framework designed to decouple initial grasp from complex interaction execution. First, Foundation-Correspondence-Guided Grasp module leverages structural priors to align the initial contact state, thereby significantly reducing the pose uncertainty of grasping. Subsequently, Multi-Foundation-Model Diffusion Policy (MFMDP) routes objects to category-specialized foundation models, integrating fine-grained geometric information with highly-variable part features via a dual-stream cross-attention mechanism. Experimental evaluations demonstrate that HeteroGenManip achieves robust intra-category shape and pose generalization. The framework achieves an average 31% performance improvement in simulation tasks with broad type setting, alongside a 36.7% gain across four real-world tasks with different interaction types.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。