arXiv:2508.11898cs.RO2025-08被引 4

用图像生成统一鸟瞰图,让机器人抓取更通用。

OmniD: Generalizable Robot Manipulation Policy via Image-Based BEV Representation

  • 多视角图像融合生成统一鸟瞰表征
  • 在分布外和少样本场景下提升84%性能
  • 适合需要强泛化的机器人抓取任务

视觉-运动策略容易过拟合于训练数据,如固定相机位置和背景,导致在分布内表现良好但在分布外泛化能力差。现有方法也难以有效融合多视角信息以生成三维表示。为此,我们提出Omni-Vision Diffusion Policy(OmniD),一种将多视角图像观测融合为统一鸟瞰图(BEV)表示的框架。引入基于可变形注意力的Omni-Feature Generator(OFG),选择性提取任务相关特征,同时抑制视角特异性噪声和背景干扰。OmniD在分布内、分布外及少样本实验中分别相比最优基线模型平均提升11%、17%和84%。训练代码与仿真基准已公开:https://github.com/1mather/omnid.git

原文摘要 · Abstract (English)

The visuomotor policy can easily overfit to its training datasets, such as fixed camera positions and backgrounds. This overfitting makes the policy perform well in the in-distribution scenarios but underperform in the out-of-distribution generalization. Additionally, the existing methods also have difficulty fusing multi-view information to generate an effective 3D representation. To tackle these issues, we propose Omni-Vision Diffusion Policy (OmniD), a multi-view fusion framework that synthesizes image observations into a unified bird's-eye view (BEV) representation. We introduce a deformable attention-based Omni-Feature Generator (OFG) to selectively abstract task-relevant features while suppressing view-specific noise and background distractions. OmniD achieves 11\%, 17\%, and 84\% average improvement over the best baseline model for in-distribution, out-of-distribution, and few-shot experiments, respectively. Training code and simulation benchmark are available: https://github.com/1mather/omnid.git

机器人抓取多视角融合泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。