arXiv:2509.15717cs.RO2025-09被引 4

让机器人在推理时‘想象’手部视角,提升抓取精度。

Imagination at Inference: Synthesizing In-Hand Views for Robust Visuomotor Policy Inference

  • 用微调的扩散模型从视角推断手部图像。
  • 真实机器人抓草莓任务中性能接近带专用摄像头方案。
  • 适合硬件受限的机器人部署,无需额外摄像头。

不同视角的视觉观测对机器人操作中的视觉-运动策略性能有显著影响,其中自参考(手部)视角常提供关键控制信息。然而,在某些应用中,为机器人配备专用手部摄像头可能受硬件限制、系统复杂性及成本制约。本文提出赋予机器人‘想象力’——在推理阶段生成手部视角图像。通过微调扩散模型(零样本视图合成,ZeroNVS),结合代理与手部摄像头间的相对位姿,实现新型视图合成(NVS)。我们基于LoRA方法对预训练模型进行领域适配,在仿真基准(RoboMimic 和 MimicGen)及真实世界实验(使用Unitree Z1机械臂完成草莓采摘任务)中验证该方法。结果表明,合成的手部视角显著提升策略推理效果,有效弥补无真实手部摄像头导致的性能下降。本方法为部署鲁棒视觉-运动策略提供了可扩展且轻量级的解决方案,凸显了具身智能体中想象式视觉推理的潜力。

原文摘要 · Abstract (English)

Visual observations from different viewpoints can significantly influence the performance of visuomotor policies in robotic manipulation. Among these, egocentric (in-hand) views often provide crucial information for precise control. However, in some applications, equipping robots with dedicated in-hand cameras may pose challenges due to hardware constraints, system complexity, and cost. In this work, we propose to endow robots with imaginative perception - enabling them to 'imagine' in-hand observations from agent views at inference time. We achieve this via novel view synthesis (NVS), leveraging a fine-tuned diffusion model conditioned on the relative pose between the agent and in-hand views cameras. Specifically, we apply LoRA-based fine-tuning to adapt a pretrained NVS model (ZeroNVS) to the robotic manipulation domain. We evaluate our approach on both simulation benchmarks (RoboMimic and MimicGen) and real-world experiments using a Unitree Z1 robotic arm for a strawberry picking task. Results show that synthesized in-hand views significantly enhance policy inference, effectively recovering the performance drop caused by the absence of real in-hand cameras. Our method offers a scalable and hardware-light solution for deploying robust visuomotor policies, highlighting the potential of imaginative visual reasoning in embodied agents.

机器人视觉推理扩散模型虚实融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。