arXiv:2603.19616cs.CV2026-03被引 1

单张立体图像端到端重建物体,精度高且省时

UniPR: Unified Object-level Real-to-Sim Perception and Reconstruction from a Single Stereo Pair

  • 直接处理单对立体图像,融合几何约束解尺度歧义
  • 一次前向传播完成全场景物体重建,保持真实物理比例
  • 适合机器人感知与仿真迁移任务,支持大规模物体

从图像中感知和重建物体是实转虚任务的关键,广泛应用于机器人领域。现有方法依赖检测、分割、形状重建和位姿估计等多个子模块组成的流水线,但存在效率低、误差累积的问题,因各阶段仅利用局部信息而忽略全局上下文。为此,我们提出UniPR,首个端到端的对象级实转虚感知与重建框架。该框架直接处理单对立体图像,利用几何约束解决尺度模糊问题。引入姿态感知的形状表示,避免依赖类别特异的规范定义,并弥合重建与位姿估计之间的差距。此外,构建了包含超过6,300个物体的大词汇量立体数据集LVS6D,以推动该领域的规模化研究。大量实验表明,UniPR可在一次前向传播中并行重建场景内所有物体,显著提升效率,同时在多种物体类型间保持真实物理比例,展现出在实际机器人应用中的巨大潜力。

原文摘要 · Abstract (English)

Perceiving and reconstructing objects from images are critical for real-to-sim transfer tasks, which are widely used in the robotics community. Existing methods rely on multiple submodules such as detection, segmentation, shape reconstruction, and pose estimation to complete the pipeline. However, such modular pipelines suffer from inefficiency and cumulative error, as each stage operates on only partial or locally refined information while discarding global context. To address these limitations, we propose UniPR, the first end-to-end object-level real-to-sim perception and reconstruction framework. Operating directly on a single stereo image pair, UniPR leverages geometric constraints to resolve the scale ambiguity. We introduce Pose-Aware Shape Representation to eliminate the need for per-category canonical definitions and to bridge the gap between reconstruction and pose estimation tasks. Furthermore, we construct a large-vocabulary stereo dataset, LVS6D, comprising over 6,300 objects, to facilitate large-scale research in this area. Extensive experiments demonstrate that UniPR reconstructs all objects in a scene in parallel within a single forward pass, achieving significant efficiency gains and preserves true physical proportions across diverse object types, highlighting its potential for practical robotic applications.

三维重建立体视觉机器人感知端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。