让机器人更懂目标地规划动作和观察,提升操作准确性。
H-GAR: A Hierarchical Interaction Framework via Goal-Driven Observation-Action Refinement for Robotic Manipulation
- 分层设计:先粗后细,用目标引导动作与观察协同生成。
- 真实场景测试表现领先,动作与观察更连贯一致。
- 适合研究机器人决策、具身智能的开发者参考。
统一的视频与动作预测模型在机器人操作中潜力巨大,因未来观察可提供规划上下文,而动作揭示交互如何改变环境。然而,现有方法通常以单一方式处理观察与动作生成,缺乏目标导向性,常导致语义错配和行为不连贯。为此,我们提出H-GAR:一种基于目标驱动的观察-动作精炼分层框架。该框架首先生成目标观察与粗粒度动作草图,勾勒出达成目标的高层路径。为在目标观察指导下实现动作与观察的显式交互,我们设计两个协同模块:(1) 目标条件观察合成器(GOS),根据粗动作与预测目标观察生成中间观察;(2) 交互感知动作精炼器(IAAR),利用中间观察反馈与历史动作记忆库,将粗动作细化为符合目标且时序一致的细粒度动作。通过在粗到细过程中融合目标锚定与显式交互,H-GAR显著提升操作精度。大量仿真与真实机器人任务实验表明,H-GAR达到当前最优性能。
原文摘要 · Abstract (English)
Unified video and action prediction models hold great potential for robotic manipulation, as future observations offer contextual cues for planning, while actions reveal how interactions shape the environment. However, most existing approaches treat observation and action generation in a monolithic and goal-agnostic manner, often leading to semantically misaligned predictions and incoherent behaviors. To this end, we propose H-GAR, a Hierarchical interaction framework via Goal-driven observation-Action Refinement.To anchor prediction to the task objective, H-GAR first produces a goal observation and a coarse action sketch that outline a high-level route toward the goal. To enable explicit interaction between observation and action under the guidance of the goal observation for more coherent decision-making, we devise two synergistic modules. (1) Goal-Conditioned Observation Synthesizer (GOS) synthesizes intermediate observations based on the coarse-grained actions and the predicted goal observation. (2) Interaction-Aware Action Refiner (IAAR) refines coarse actions into fine-grained, goal-consistent actions by leveraging feedback from the intermediate observations and a Historical Action Memory Bank that encodes prior actions to ensure temporal consistency. By integrating goal grounding with explicit action-observation interaction in a coarse-to-fine manner, H-GAR enables more accurate manipulation. Extensive experiments on both simulation and real-world robotic manipulation tasks demonstrate that H-GAR achieves state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。