arXiv:2609.07006cs.CVcs.RO2026-09

用生成目标图像提升机器人操控模型效率与鲁棒性

GIFT: Goal-Injected Fine-Tuning for Efficient Manipulation Policy Adaptation

论文配图:GIFT: Goal-Injected Fine-Tuning for Efficient Manipulation Policy Adaptation
图 1 · 摘自论文原文
  • 通过零初始化卷积逐步注入生成的目标图像特征
  • 单轮微调即在多个任务上提升4.7%~13.4%
  • 适合需快速适配新任务的机器人应用

相较于仅依赖初始观测和语言指令,利用生成模型预测目标图像作为高层视觉引导,可显著提升视觉-语言-动作(VLA)模型的鲁棒性。然而,由于训练成本高,多数现有基础模型未系统引入目标图像条件。为此,我们提出轻量级高效微调框架GIFT,将生成的目标图像无缝集成至多个代表性预训练VLA模型中。方法通过零初始化卷积将目标图像特征引入观测,参数从零渐进增长,避免噪声干扰预训练策略。随着训练推进,目标信息逐步融入,实现高效目标理解且不破坏模型稳定性。此外,我们设计了一种优化图像编辑方法,从初始观测和任务指令生成语义与视觉一致的目标图像。实验表明,目标感知的VLA模型在多个任务上表现显著提升:仅需单轮微调,GIFT在两个SIMPLER设置上分别优于基线6.0%和13.4%,在LIBERO上提升4.7%,验证了其高效性与有效性。

原文摘要 · Abstract (English)

Compared with relying solely on initial observations and language instructions, predicting goal images with generative models as high-level visual guidance can significantly enhance the robustness of Vision-Language-Action (VLA) models. However, most existing foundation models have not systematically incorporated goal image conditioning due to the high computational training cost. To this end, we propose Goal-Injected Fine-Tuning (GIFT), a lightweight and efficient fine-tuning framework that seamlessly integrates generated goal images into multiple representative pretrained VLA models. Our approach introduces goal image features into observations via a zero-initialized convolution which progressively grows parameters from zero and prevents harmful noise from disrupting the pretrained policy during fine-tuning. As training proceeds, goal information is gradually incorporated, enabling efficient goal understanding without disrupting model stability. We further introduce a refined image editing method to generate semantically and visually consistent goal images from initial observations and task instructions. Experiments show that goal-aware VLA models achieve substantial performance gains across tasks: with only a single epoch of fine-tuning, GIFT outperforms the base model by 6.0% and 13.4% on two SIMPLER settings, and by 4.7% on LIBERO, demonstrating both efficiency and effectiveness.

机器人操控视觉语言微调生成目标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。