arXiv:2604.19257cs.CV2026-04中稿 · CVPR

从真实道路图像生成可直接用于仿真环境的3D车辆模型

Unposed-to-3D: Learning Simulation-Ready Vehicles from Real-World Images

论文配图:Unposed-to-3D: Learning Simulation-Ready Vehicles from Real-World Images
图 1 · 摘自论文原文
  • 仅用无姿态图像训练,通过自监督光度反馈学习3D结构
  • 重建出姿态一致、尺度真实且与场景光照匹配的3D车辆
  • 适合自动驾驶仿真与数字孪生场景的高质量资产构建

创建真实且可直接用于仿真的3D资产对自动驾驶研究和虚拟环境构建至关重要。现有3D车辆生成方法多基于合成数据训练,存在与真实分布差异大、模型姿态任意、尺度未定义等问题,导致在驾驶场景中视觉不一致。本文提出Unposed-to-3D框架,仅使用真实道路图像进行监督,分两阶段重建3D车辆:第一阶段用带姿态图像与已知相机参数训练图像到3D重建网络;第二阶段移除相机监督,引入相机预测头,从无姿态图像中估计相机参数,并通过可微渲染提供自监督光度反馈,使模型仅凭无姿态图像学习3D几何。为确保仿真可用性,进一步引入尺度感知模块预测真实尺寸,以及和谐化模块适配生成车辆至目标驾驶场景的光照与外观。大量实验表明,Unposed-to-3D能从真实图像中有效重建出姿态一致、视觉协调的3D车辆模型,为驾驶场景仿真与数字孪生环境提供可扩展的高质量资产生成路径。

原文摘要 · Abstract (English)

Creating realistic and simulation-ready 3D assets is crucial for autonomous driving research and virtual environment construction. However, existing 3D vehicle generation methods are often trained on synthetic data with significant domain gaps from real-world distributions. The generated models often exhibit arbitrary poses and undefined scales, resulting in poor visual consistency when integrated into driving scenes. In this paper, we present Unposed-to-3D, a novel framework that learns to reconstruct 3D vehicles from real-world driving images using image-only supervision. Our approach consists of two stages. In the first stage, we train an image-to-3D reconstruction network using posed images with known camera parameters. In the second stage, we remove camera supervision and use a camera prediction head that directly estimates the camera parameters from unposed images. The predicted pose is then used for differentiable rendering to provide self-supervised photometric feedback, enabling the model to learn 3D geometry purely from unposed images. To ensure simulation readiness, we further introduce a scale-aware module to predict real-world size information, and a harmonization module that adapts the generated vehicles to the target driving scene with consistent lighting and appearance. Extensive experiments demonstrate that Unposed-to-3D effectively reconstructs realistic, pose-consistent, and harmonized 3D vehicle models from real-world images, providing a scalable path toward creating high-quality assets for driving scene simulation and digital twin environments.

3D生成自动驾驶自监督数字孪生

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。