从自动驾驶日志中提取完整3D资产,用于仿真环境构建
Asset Harvester: Extracting 3D Assets from Autonomous Driving Logs for Simulation

- 构建端到端系统,将稀疏观测图像转为完整可交互3D模型
- 在真实驾驶数据上实现多视角生成与3D高斯提升,支持新视角合成
- 适合需要大规模仿真资产的自动驾驶研发团队使用
闭环仿真在自动驾驶开发中至关重要,可实现规模化测试、训练与安全验证。神经场景重建虽能将驾驶日志转为交互式3D环境,但无法生成可用于代理操控和大视角新视图合成的完整3D物体资产。为此,我们提出Asset Harvester,一个从真实驾驶日志中稀疏、野生对象观测转换为完整仿真可用资产的图像到3D模型端到端管道。系统结合大规模对象中心训练样本采集、跨异构传感器的几何感知预处理,以及耦合稀疏视图条件多视角生成与3D高斯提升的鲁棒训练策略。其中,SparseViewDiT专门应对有限视角等现实数据挑战。结合混合数据采集、增强与自蒸馏,该系统可规模化将稀疏自动驾驶物体观测转化为可复用3D资产。
原文摘要 · Abstract (English)
Closed-loop simulation is a core component of autonomous vehicle (AV) development, enabling scalable testing, training, and safety validation before real-world deployment. Neural scene reconstruction converts driving logs into interactive 3D environments for simulation, but it does not produce complete 3D object assets required for agent manipulation and large-viewpoint novel-view synthesis. To address this challenge, we present Asset Harvester, an image-to-3D model and end-to-end pipeline that converts sparse, in-the-wild object observations from real driving logs into complete, simulation-ready assets. Rather than relying on a single model component, we developed a system-level design for real-world AV data that combines large-scale curation of object-centric training tuples, geometry-aware preprocessing across heterogeneous sensors, and a robust training recipe that couples sparse-view-conditioned multiview generation with 3D Gaussian lifting. Within this system, SparseViewDiT is explicitly designed to address limited-angle views and other real-world data challenges. Together with hybrid data curation, augmentation, and self-distillation, this system enables scalable conversion of sparse AV object observations into reusable 3D assets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。