arXiv:2604.08475cs.CV2026-04

用图像编辑生成3D空间关系,让机器人在未知环境里精准操作

LAMP: Lift Image-Editing as General 3D Priors for Open-world Manipulation

论文配图:LAMP: Lift Image-Editing as General 3D Priors for Open-world Manipulation
图 1 · 摘自论文原文
  • 将图像编辑的2D空间线索转化为连续3D变换表示
  • 在开放世界任务中实现高精度3D操作与零样本泛化
  • 适合需要灵活适应新环境的机器人操作研究者

开放世界中的类人泛化仍是机器人操作的核心挑战。现有基于学习的方法(如强化学习、模仿学习、视觉语言动作模型)常难以应对新任务和未见环境。另一条有前景的方向是探索能捕捉细粒度空间与几何关系的通用表征。尽管大语言模型和视觉语言模型具备强大的语义推理能力,但其有限的3D感知限制了在精细操作中的应用。为此,我们提出LAMP,通过将图像编辑作为3D先验,提取物体间连续的、几何感知的3D变换表示。核心洞察在于:图像编辑天然蕴含丰富的2D空间线索,将其提升为3D变换可提供精确的细粒度引导。大量实验表明,LAMP能生成准确的3D变换,并在开放世界操作中实现强零样本泛化。项目页:https://zju3dv.github.io/LAMP/

原文摘要 · Abstract (English)

Human-like generalization in open-world remains a fundamental challenge for robotic manipulation. Existing learning-based methods, including reinforcement learning, imitation learning, and vision-language-action-models (VLAs), often struggle with novel tasks and unseen environments. Another promising direction is to explore generalizable representations that capture fine-grained spatial and geometric relations for open-world manipulation. While large-language-model (LLMs) and vision-language-model (VLMs) provide strong semantic reasoning based on language or annotated 2D representations, their limited 3D awareness restricts their applicability to fine-grained manipulation. To address this, we propose LAMP, which lifts image-editing as 3D priors to extract inter-object 3D transformations as continuous, geometry-aware representations. Our key insight is that image-editing inherently encodes rich 2D spatial cues, and lifting these implicit cues into 3D transformations provides fine-grained and accurate guidance for open-world manipulation. Extensive experiments demonstrate that \codename delivers precise 3D transformations and achieves strong zero-shot generalization in open-world manipulation. Project page: https://zju3dv.github.io/LAMP/.

机器人操作3D理解零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。