用流模型生成具语义的双手操作动作,提升机器人灵巧操作成功率。
FlowHOI: Flow-based Semantics-Grounded Generation of Hand-Object Interactions for Dexterous Robot Manipulation
- 分两阶段流匹配生成手物交互序列,解耦抓取与操作任务。
- 在GRAB和HOT3D上动作识别准确率最高,物理仿真成功率提升1.7倍。
- 支持真实机器人执行,推理速度比扩散模型快40倍。
当前视觉-语言-动作模型虽能生成合理末端运动,但在长时程、高接触任务中常因未显式建模手物交互(HOI)结构而失败。我们提出FlowHOI,一种两阶段流匹配框架,基于自中心观测、语言指令和3D高斯点云重建(3DGS),生成语义对齐、时间连贯的手姿、物姿及手物接触状态序列。该方法将几何导向的抓取与语义导向的操作解耦,后者以紧凑的3D场景令牌为条件,并引入运动-文本对齐损失,使生成交互同时契合物理布局与语言指令。针对高质量HOI标注数据稀缺问题,我们设计了一套重建流程,从大规模自中心视频中恢复对齐的手物轨迹与网格,构建用于鲁棒生成的HOI先验。在GRAB和HOT3D基准上,FlowHOI达到最高动作识别准确率,物理仿真成功率达最强扩散基线的1.7倍,且推理速度提升40倍。我们进一步在四类灵巧操作任务上实现真实机器人执行,验证了生成的HOI表示可直接适配真实机器人执行管道。
原文摘要 · Abstract (English)
Recent vision-language-action (VLA) models can generate plausible end-effector motions, yet they often fail in long-horizon, contact-rich tasks because the underlying hand-object interaction (HOI) structure is not explicitly represented. An embodiment-agnostic interaction representation that captures this structure would make manipulation behaviors easier to validate and transfer across robots. We propose FlowHOI, a two-stage flow-matching framework that generates semantically grounded, temporally coherent HOI sequences, comprising hand poses, object poses, and hand-object contact states, conditioned on an egocentric observation, a language instruction, and a 3D Gaussian splatting (3DGS) scene reconstruction. We decouple geometry-centric grasping from semantics-centric manipulation, conditioning the latter on compact 3D scene tokens and employing a motion-text alignment loss to semantically ground the generated interactions in both the physical scene layout and the language instruction. To address the scarcity of high-fidelity HOI supervision, we introduce a reconstruction pipeline that recovers aligned hand-object trajectories and meshes from large-scale egocentric videos, yielding an HOI prior for robust generation. Across the GRAB and HOT3D benchmarks, FlowHOI achieves the highest action recognition accuracy and a 1.7$\times$ higher physics simulation success rate than the strongest diffusion-based baseline, while delivering a 40$\times$ inference speedup. We further demonstrate real-robot execution on four dexterous manipulation tasks, illustrating the feasibility of retargeting generated HOI representations to real-robot execution pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。