用扩散模型让自动驾驶仿真到现实的差距缩小40%以上
Sim2Real Diffusion: Leveraging Foundation Vision Language Models for Adaptive Automated Driving
- 基于条件扩散生成跨域适应表征,支持多模态提示输入
- 在少样本条件下仍保持鲁棒性能,实测模拟到现实的感知差距降低超40%
- 可灵活适配不同场景(如天气、时段),适合自动驾驶系统验证
基于仿真的自动驾驶设计、优化与验证已证明对系统改进至关重要。然而,其最终有效性取决于能否成功从仿真环境过渡到真实世界(sim2real)。现有方法难以兼顾:(i) 条件化领域自适应,(ii) 有限样本下的鲁棒性能,(iii) 多种领域表示的模块化处理,以及 (iv) 实时性要求。为此,我们提出一种统一框架,通过条件潜在扩散学习跨域自适应表征,实现可迁移的自动驾驶 sim2real 转移。该框架支持:(i) 交替使用基础视觉语言模型,(ii) 少样本微调流程,(iii) 文本与图像提示以映射源域与目标域。同时具备在时间、天气、季节及运行设计域等参数空间中生成多样化高质量样本的能力。我们系统分析了该框架,并在性能基准和消融实验中报告结果。此外,通过行为克隆案例研究验证其在自动驾驶中的实用性。实验表明,所提框架可将感知层面的 sim2real 差距缩小超过40%。
原文摘要 · Abstract (English)
Simulation-based design, optimization, and validation of autonomous vehicles have proven to be crucial for their improvement over the years. Nevertheless, the ultimate measure of effectiveness is their successful transition from simulation to reality (sim2real). However, existing sim2real transfer methods struggle to address the autonomy-oriented requirements of balancing: (i) conditioned domain adaptation, (ii) robust performance with limited examples, (iii) modularity in handling multiple domain representations, and (iv) real-time performance. To alleviate these pain points, we present a unified framework for learning cross-domain adaptive representations through conditional latent diffusion for sim2real transferable automated driving. Our framework offers options to leverage: (i) alternate foundation models, (ii) a few-shot fine-tuning pipeline, and (iii) textual as well as image prompts for mapping across given source and target domains. It is also capable of generating diverse high-quality samples when diffusing across parameter spaces such as times of day, weather conditions, seasons, and operational design domains. We systematically analyze the presented framework and report our findings in terms of performance benchmarks and ablation studies. Additionally, we demonstrate its serviceability for autonomous driving using behavioral cloning case studies. Our experiments indicate that the proposed framework is capable of bridging the perceptual sim2real gap by over 40%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。