无需物体几何信息,从手物交互图像生成多模态抓取姿态
HOGraspFlow: Taxonomy-Aware Hand-Object Retargeting for Multi-Modal SE(3) Grasp Generation
- 基于视觉语义、接触重建与分类先验,用去噪流匹配生成抓取位姿
- 真实场景下平均成功率超83%,无需目标物体几何数据
- 适合无实物模型的机器人抓取应用,尤其擅长从人演示中学习
我们提出HOGraspFlow,一种以功能为导向的方法,将单个带手物交互的RGB图像转换为无需目标物体显式几何先验的多模态可执行平行指抓取。基于手部重建和视觉基础模型,通过条件化于三类互补线索——RGB基础特征(视觉语义)、手物接触重建以及抓取类型分类感知先验——利用去噪流匹配(FM)合成SE(3)抓取位姿。该方法在无显式手物接触输入或物体几何信息条件下仍保持高保真抓取生成,同时具备强接触与分类识别能力。控制对比实验表明,HOGraspFlow持续优于基于扩散模型的变体(HOGraspDiff),在SE(3)空间中实现更高分布保真度与更稳定优化。真实世界实验验证了其从人类示范中实现可靠、无对象依赖的抓取生成,平均成功率达83%以上。
原文摘要 · Abstract (English)
We propose Hand-Object\emph{(HO)GraspFlow}, an affordance-centric approach that retargets a single RGB with hand-object interaction (HOI) into multi-modal executable parallel jaw grasps without explicit geometric priors on target objects. Building on foundation models for hand reconstruction and vision, we synthesize $SE(3)$ grasp poses with denoising flow matching (FM), conditioned on the following three complementary cues: RGB foundation features as visual semantics, HOI contact reconstruction, and taxonomy-aware prior on grasp types. Our approach demonstrates high fidelity in grasp synthesis without explicit HOI contact input or object geometry, while maintaining strong contact and taxonomy recognition. Another controlled comparison shows that \emph{HOGraspFlow} consistently outperforms diffusion-based variants (\emph{HOGraspDiff}), achieving high distributional fidelity and more stable optimization in $SE(3)$. We demonstrate a reliable, object-agnostic grasp synthesis from human demonstrations in real-world experiments, where an average success rate of over $83\%$ is achieved. Code: https://github.com/YitianShi/HOGraspFlow
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。