arXiv:2511.22777cs.ROcs.AI2025-11被引 2

用AI生成新场景提升机器人抓取鲁棒性,无需额外数据收集。

Improving Robotic Manipulation Robustness via NICE Scene Surgery

  • 通过图像生成与语言模型修改场景,保留空间关系并增强视觉多样性。
  • 在杂乱环境中物体识别准确率提升超20%,抓取成功率平均提高11%。
  • 不需新数据或训练,适合已有机器人数据集快速升级应用。

真实场景中视觉干扰物会显著降低机器人抓取的性能与安全性。本文提出自然化修复增强框架NICE,通过现有示范数据构建新场景以缩小分布外(OOD)差距。利用图像生成模型和大语言模型,NICE执行对象替换、风格重制及干扰物移除三种编辑操作,保持目标物体空间关系完整与动作标签一致。相比以往方法,NICE无需额外机器人数据采集、仿真环境或定制模型训练,可直接应用于现有数据集。在真实场景中验证了其生成逼真场景的能力。下游任务中,使用NICE数据微调视觉-语言模型(VLM)进行空间可操作性预测,以及视觉-语言-动作(VLA)策略进行物体操作。评估显示,NICE有效缩小了OOD差距,使复杂杂乱场景下可操作性预测准确率提升超过20%;抓取任务在不同数量干扰物环境下平均成功率提高11%。同时,目标混淆率降低6%,碰撞率下降7%,显著提升视觉鲁棒性与安全性。

原文摘要 · Abstract (English)

Learning robust visuomotor policies for robotic manipulation remains a challenge in real-world settings, where visual distractors can significantly degrade performance and safety. In this work, we propose an effective and scalable framework, Naturalistic Inpainting for Context Enhancement (NICE). Our method minimizes out-of-distribution (OOD) gap in imitation learning by increasing visual diversity through construction of new experiences using existing demonstrations. By utilizing image generative frameworks and large language models, NICE performs three editing operations, object replacement, restyling, and removal of distracting (non-target) objects. These changes preserve spatial relationships without obstructing target objects and maintain action-label consistency. Unlike previous approaches, NICE requires no additional robot data collection, simulator access, or custom model training, making it readily applicable to existing robotic datasets. Using real-world scenes, we showcase the capability of our framework in producing photo-realistic scene enhancement. For downstream tasks, we use NICE data to finetune a vision-language model (VLM) for spatial affordance prediction and a vision-language-action (VLA) policy for object manipulation. Our evaluations show that NICE successfully minimizes OOD gaps, resulting in over 20% improvement in accuracy for affordance prediction in highly cluttered scenes. For manipulation tasks, success rate increases on average by 11% when testing in environments populated with distractors in different quantities. Furthermore, we show that our method improves visual robustness, lowering target confusion by 6%, and enhances safety by reducing collision rate by 7%.

机器人操控视觉鲁棒性生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。