arXiv:2606.22987cs.CVcs.RO2026-06

研究机器人相机旋转对单视角网格重建的影响,发现现有方法泛化能力差。

Can Single-View Mesh Reconstruction Generalize to Robot Camera Rotation?

论文配图:Can Single-View Mesh Reconstruction Generalize to Robot Camera Rotation?
图 1 · 摘自论文原文
  • 通过控制旋转轴扫面测试不同视角下的重建误差
  • 相机旋转导致深度估计失真、布局漂移和碰撞穿透,但网格预测较稳定
  • 引入重力感知优化,显著提升单阶段布局预测鲁棒性

单视角网格重建可从单一观测中预测物体网格与空间布局,适用于机器人快速空间推理和真实世界到数字孪生的映射。然而,机器人安装的相机在操作和导航中会自然旋转,而现有单视角重建模型依赖视图相关先验,对分布外的相机旋转泛化能力差。这种旋转会导致三维不一致、错误布局和违反物理约束,但该失效模式尚未被充分评估。本文提出一种评估协议,通过控制轴向滚转、俯仰和偏航的扫面,追踪代表性SAM3D风格流程中单目深度估计(MDE)、标准物体网格、相机空间布局及物理合理性中的误差。在Aria Digital Twin数据集和真实Franka腕部相机序列上,相机旋转引发MDE失真、布局漂移和碰撞穿透,而标准网格预测相对稳定。两阶段SAM3D+FoundationPose流程比单阶段前馈布局预测更鲁棒,且我们的重力感知优化使单阶段基于ICP的布局方向误差降低47.1%。评估揭示当前单视角网格重建方法对机器人相机旋转泛化不佳,提示显式重力线索对可靠机器人单视角重建至关重要。

原文摘要 · Abstract (English)

Single-view mesh reconstruction predicts object meshes and spatial layouts from a single observation, making it attractive for fast robot spatial reasoning and real-to-sim digital twins. However, robot-mounted cameras naturally rotate during manipulation and navigation, while learned single-view reconstruction models often rely on view-dependent priors and may generalize poorly to out-of-distribution camera rotations. Such rotations can introduce 3D inconsistencies, incorrect layouts, and violations of physical constraints, but this failure mode remains under-evaluated. We introduce an evaluation protocol with controlled axis-wise roll, pitch, and yaw sweeps to trace errors in monocular depth estimation (MDE), canonical object meshes, camera-space layout, and physical plausibility within a representative SAM3D-style pipeline. On the Aria Digital Twin dataset and a real Franka wrist-camera sequence, camera rotations induce MDE distortion, layout drift, and collision penetration, while canonical mesh predictions remain relatively stable. A two-stage SAM3D+FoundationPose pipeline is more robust than one-stage feed-forward layout prediction, and our Gravity-Aware Refinement reduces one-stage pairwise ICP-based layout-orientation error by 47.1$\%$. Our evaluation reveals that current single-view mesh reconstruction methods generalize poorly to robot camera rotation, and suggests that explicit gravity cues are important for reliable robotic single-view mesh reconstruction.

单视角重建机器人视觉重力感知数字孪生

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。