统一框架实现多视角3D重建,无需重新训练即可获度量尺度。
UniScale: Unified Scale-Aware 3D Reconstruction for Multi-View Understanding via Prior Injection for Robotic Perception
- 模块化设计注入几何先验,联合估计相机参数与场景尺度。
- 在已知相机内参时精度提升,结合位姿信息进一步优化结果。
- 适合资源受限机器人团队,兼容预训练模型快速部署。
我们提出UniScale,一种面向机器人感知的统一、尺度感知的多视角3D重建框架,通过模块化语义引导设计灵活融合几何先验。在视觉导航中,从原始图像序列准确提取环境结构对下游任务至关重要。UniScale采用单次前向网络,联合估计相机内参与外参、尺度不变深度图与点云,并恢复场景的度量尺度,同时可选地引入辅助几何先验。通过结合全局上下文推理与相机感知特征表示,实现场景度量尺度重建。在相机内参已知的机器人场景中,可轻松融入以提升性能;当相机位姿可用时,进一步获得增益。该协同设计使单一统一模型实现鲁棒的度量感知3D重建。重要的是,UniScale无需从头训练,利用现有模型中的世界先验而无需几何编码策略,特别适用于资源受限的机器人团队。我们在多个基准上评估了UniScale,证明其在多样化环境中的强泛化能力和一致性能。代码将在论文接收后公开。
原文摘要 · Abstract (English)
We present UniScale, a unified, scale-aware multi-view 3D reconstruction framework for robotic applications that flexibly integrates geometric priors through a modular, semantically informed design. In vision-based robotic navigation, the accurate extraction of environmental structure from raw image sequences is critical for downstream tasks. UniScale addresses this challenge with a single feed-forward network that jointly estimates camera intrinsics and extrinsics, scale-invariant depth and point maps, and the metric scale of a scene from multi-view images, while optionally incorporating auxiliary geometric priors when available. By combining global contextual reasoning with camera-aware feature representations, UniScale is able to recover the metric-scale of the scene. In robotic settings where camera intrinsics are known, they can be easily incorporated to improve performance, with additional gains obtained when camera poses are also available. This co-design enables robust, metric-aware 3D reconstruction within a single unified model. Importantly, UniScale does not require training from scratch, and leverages world priors exhibited in pre-existing models without geometric encoding strategies, making it particularly suitable for resource-constrained robotic teams. We evaluate UniScale on multiple benchmarks, demonstrating strong generalization and consistent performance across diverse environments. We will release our implementation upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。