arXiv:2508.01442cs.RO2025-08被引 8

用物理光照生成技术提升机器人在不同光线下的操作能力

Physically-based Lighting Generation for Robotic Manipulation

  • 通过逆向渲染提取物体几何与材质,生成新光照下的视觉变化
  • 在六种未知光照条件下使模仿学习策略性能提升38.75%
  • 适用于机器人训练数据增强与真实场景模拟

本文提出首个基于物理的逆向渲染框架,用于在现有真实世界人类示范视频上生成新光照条件下的机器人操作序列。具体而言,逆向渲染将每个示范的第一帧分解为几何(法线、深度)和材质(反照率、粗糙度、金属度)属性,并据此渲染不同光源下的外观变化。为提升效率并保证序列一致性,我们对机器人执行视频微调Stable Video Diffusion实现时序光照传播。评估显示,该框架在六种未见真实光照条件下,使模仿学习策略性能提升38.75%。同时通过消融实验验证各模块有效性,并展示三项下游应用:背景生成、物体纹理生成与干扰物定位。代码将公开。

原文摘要 · Abstract (English)

In this paper, we propose the first framework that leverages physically-based inverse rendering for novel lighting generation on existing real-world human demonstrations of robotic manipulation tasks. Specifically, inverse rendering decomposes the first frame in each demonstration into geometric (surface normal, depth) and material (albedo, roughness, metallic) properties, which are then used to render appearance changes under different lighting sources. To improve efficiency and maintain consistency across each generated sequence, we fine-tune Stable Video Diffusion on robot execution videos for temporal lighting propagation. We evaluate our framework by measuring the visual quality of the generated sequences, assessing its effectiveness in improving the imitation learning policy performance (38.75\%) under six unseen real-world lighting conditions, and conduct ablation studies on individual modules of the proposed framework. We further showcase three downstream applications enabled by the proposed framework: background generation, object texture generation and distractor positioning. The code for the framework will be made publicly available.

光照生成机器人操作逆向渲染模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。