arXiv:2609.04193cs.RO2026-09

通过结构化监督让机器人视觉特征更懂动作需求,提升复杂操作能力。

GIFT: Guided Intermediate Feature Training via Action-Oriented Structural Supervision for Robotic Manipulation

论文配图:GIFT: Guided Intermediate Feature Training via Action-Oriented Structural Supervision for Robotic Manipulation
图 1 · 摘自论文原文
  • 用几何、功能和目标区域三重约束指导中间特征学习
  • 在LIBERO-Plus和RoboCasa上分别提升12.6%和9.0%以上准确率
  • 特别适合需要高精度操控的复杂物体任务

视觉语言预训练和预测世界建模为机器人策略提供了丰富的语义与动态视觉特征,但其原始的动作与视觉预测目标可能忽略关键物理与任务结构,同时保留无关视觉冗余。我们称这种视觉丰富性与控制效用之间的差距为动作充分性缺口。本文提出GIFT(Guided Intermediate Feature Training),通过几何对齐、功能预测和目标区域重建,将运动可行性、指令相关实体功能、任务相关区域目标这三类控制相关结构转化为训练时的约束,引导中间特征学习。在视觉-语言-动作(VLA)策略、直接动作世界-动作模型(WAM)和逆动力学WAM中均成功应用,保持原模型动作形式。零样本迁移至LIBERO-Plus时,GIFT-VLA、GIFT-WAM-Fast、GIFT-WAM-IDM分别比StarVLA-OFT、Fast-WAM、Fast-WAM-IDM高出4.6、12.6、5.2点,达到79.6%、72.6%、87.8%;在RoboCasa上,三者分别达61.4%、83.6%、82.3%,优于基线12.6、9.0、8.4点。结果表明,学习功能性结构化中间特征是一种可跨模型复用的原则,尤其在可动物体任务和未见视觉/空间扰动下的高精度真实操控中表现突出。

原文摘要 · Abstract (English)

Vision-language pre-training and predictive world modeling provide robot policies with rich semantic and dynamic visual features, but their native action and visual-prediction objectives may omit critical physical and task structure while retaining control-irrelevant visual redundancy. We call this mismatch between visual richness and control utility the action-sufficiency gap. We investigate whether this gap can be bridged by guiding intermediate features to preserve three control-relevant structure in robotic manipulation: geometry governing motion feasibility, affordance encoding instruction-relevant entities, and goals grounding instructions in task-relevant regions. To this end, we present GIFT (Guided Intermediate Feature Training), an architecture-flexible framework for learning intermediate features that translates these structures into training-time constraints through geometry alignment, affordance prediction, and goal-region reconstruction. We instantiate GIFT in a Vision-Language-Action (VLA) policy, a direct-action World-Action Model (WAM), and an inverse-dynamics WAM while retaining each model's action formulation. Under zero-shot transfer to LIBERO-Plus, GIFT-VLA, GIFT-WAM-Fast, and GIFT-WAM-IDM outperform StarVLA-OFT, Fast-WAM, and Fast-WAM-IDM by 4.6, 12.6, and 5.2 points, reaching 79.6%, 72.6%, and 87.8%, respectively. On RoboCasa, the three GIFT variants reach 61.4%, 83.6%, and 82.3%, outperforming their counterparts by 12.6, 9.0, and 8.4 points, respectively. Together, these results establish learning functionally structured intermediate features as a reusable principle across model-specific action formulations, with especially large gains on articulated-object tasks and high-precision real-world manipulation under unseen visual and spatial perturbations. Project page: https://openphoenix-team.github.io/GIFT-pages.

机器人操控视觉语言结构化特征零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。