用大模型让自动驾驶仿真更真实、可交互、能编辑。
When Digital Twins Meet Large Language Models: Realistic, Interactive, and Editable Simulation for Autonomous Driving
- 融合物理与数据驱动方法构建高保真数字孪生
- 实现97%结构相似度与60Hz以上实时动态模拟
- 支持自然语言编辑场景,适合自动驾驶研发人员
仿真框架是自动驾驶系统开发与验证的关键。但现有方法难以兼顾动态保真度、照片级渲染、上下文相关场景编排和实时性能。为此,我们提出统一框架,构建高保真数字孪生以加速自动驾驶研究。该框架结合物理建模与数据驱动技术,重建真实场景与资产,几何与视觉保真度达约97%,并赋予物理属性,实现超过60 Hz的实时动态模拟。通过集成大语言模型(LLM)接口,可自然语言在线编辑场景,通用性约85%,重复性达95%。此外,可选视觉语言模型(VLM)通过混合场景合成提升视觉效果,增强率约80%。
原文摘要 · Abstract (English)
Simulation frameworks have been key enablers for the development and validation of autonomous driving systems. However, existing methods struggle to comprehensively address the autonomy-oriented requirements of balancing: (i) dynamical fidelity, (ii) photorealistic rendering, (iii) context-relevant scenario orchestration, and (iv) real-time performance. To address these limitations, we present a unified framework for creating and curating high-fidelity digital twins to accelerate advancements in autonomous driving research. Our framework leverages a mix of physics-based and data-driven techniques for developing and simulating digital twins of autonomous vehicles and their operating environments. It is capable of reconstructing real-world scenes and assets with geometric and photorealistic accuracy (~97% structural similarity) and infusing them with physical properties to enable real-time (>60 Hz) dynamical simulation of the ensuing driving scenarios. Additionally, it incorporates a large language model (LLM) interface to flexibly edit the driving scenarios online via natural language prompts, with ~85% generalizability and ~95% repeatability. Finally, an optional vision language model (VLM) provides ~80% visual enhancement by blending the hybrid scene composition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。