arXiv:2604.18468cs.CVcs.AI2026-04

从自动驾驶日志中提取完整3D资产,用于仿真环境构建

Asset Harvester: Extracting 3D Assets from Autonomous Driving Logs for Simulation

论文配图:Asset Harvester: Extracting 3D Assets from Autonomous Driving Logs for Simulation
图 1 · 摘自论文原文
  • 构建端到端系统,将稀疏观测图像转为完整可交互3D模型
  • 在真实驾驶数据上实现多视角生成与3D高斯提升,支持新视角合成
  • 适合需要大规模仿真资产的自动驾驶研发团队使用

闭环仿真在自动驾驶开发中至关重要,可实现规模化测试、训练与安全验证。神经场景重建虽能将驾驶日志转为交互式3D环境,但无法生成可用于代理操控和大视角新视图合成的完整3D物体资产。为此,我们提出Asset Harvester,一个从真实驾驶日志中稀疏、野生对象观测转换为完整仿真可用资产的图像到3D模型端到端管道。系统结合大规模对象中心训练样本采集、跨异构传感器的几何感知预处理,以及耦合稀疏视图条件多视角生成与3D高斯提升的鲁棒训练策略。其中,SparseViewDiT专门应对有限视角等现实数据挑战。结合混合数据采集、增强与自蒸馏,该系统可规模化将稀疏自动驾驶物体观测转化为可复用3D资产。

原文摘要 · Abstract (English)

Closed-loop simulation is a core component of autonomous vehicle (AV) development, enabling scalable testing, training, and safety validation before real-world deployment. Neural scene reconstruction converts driving logs into interactive 3D environments for simulation, but it does not produce complete 3D object assets required for agent manipulation and large-viewpoint novel-view synthesis. To address this challenge, we present Asset Harvester, an image-to-3D model and end-to-end pipeline that converts sparse, in-the-wild object observations from real driving logs into complete, simulation-ready assets. Rather than relying on a single model component, we developed a system-level design for real-world AV data that combines large-scale curation of object-centric training tuples, geometry-aware preprocessing across heterogeneous sensors, and a robust training recipe that couples sparse-view-conditioned multiview generation with 3D Gaussian lifting. Within this system, SparseViewDiT is explicitly designed to address limited-angle views and other real-world data challenges. Together with hybrid data curation, augmentation, and self-distillation, this system enables scalable conversion of sparse AV object observations into reusable 3D assets.

3D生成自动驾驶仿真多视角重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。