arXiv:2603.07530cs.RO2026-03中稿 · IROS 2026被引 1

让机器人通过视觉推理理解任务意图,提升少样本学习效果。

ICLR: In-Context Imitation Learning with Visual Reasoning

  • 用图像空间的未来轨迹推理增强示范提示
  • 联合生成推理过程与动作,成功率显著提升
  • 适合复杂、多目标的机器人任务泛化场景

上下文模仿学习使机器人能在不额外训练的情况下,仅凭少量示范适应新任务。然而,现有方法通常仅依赖状态-动作轨迹,缺乏对任务意图的显式表示,这在复杂且模糊的任务环境中限制了性能,因为相同动作可能对应不同目标。为此,我们提出一种名为「基于视觉推理的上下文模仿学习」(ICLR)的新框架,通过在示范提示中加入结构化的视觉推理轨迹(预测未来图像空间中的机器人轨迹),增强任务理解。ICLR 还在一个统一的自回归变换器中联合学习生成推理轨迹与低层动作,使模型不仅能预测行为,还能模拟其背后的推理过程。我们在仿真和真实世界操作任务中进行了广泛评估,结果表明,相比其他上下文模仿学习方法,ICLR 在未见过的任务和新物体配置下均表现出更高的成功率与更强的泛化能力。这些结果表明,引入具身视觉推理是提升机器人上下文学习系统鲁棒性与泛化性的有效方向。

原文摘要 · Abstract (English)

In-context imitation learning enables robots to adapt to new tasks from a small number of demonstrations without additional training. However, existing approaches typically condition only on state-action trajectories and lack explicit representations of task intent. This limitation hinders performance in complex and ambiguous task settings where the same actions may be consistent with different objectives. To address this, we present In-Context Imitation Learning with Visual Reasoning (ICLR), a novel framework that augments demonstration prompts with structured visual reasoning traces representing anticipated future robot trajectories in image space. ICLR also jointly learns to generate reasoning traces and low-level actions within a unified autoregressive transformer, enabling the model to mimic not only action prediction but also the reasoning process that leads to those actions. We extensively evaluate ICLR in both simulation and real-world manipulation tasks and demonstrate consistent improvements in success rates and generalization to unseen tasks and novel object configurations compared to other in-context imitation learning methods. These results suggest that incorporating embodied visual reasoning represents a promising direction for enhancing the robustness and generalization of robotic in-context learning systems.

机器人学习视觉推理少样本学习模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。