用事故报告生成高保真自动驾驶危险场景,准确率提升182%
SG-CADVLM: A Context-Aware Decoding Powered Vision Language Model for Safety-Critical Scenario Generation
- 基于上下文感知解码,融合多模态输入生成场景
- 生成危急场景成功率88.1%,较基线提升182%
- 适合自动驾驶安全验证与仿真测试研究者
自动驾驶需在安全关键场景下进行严格测试,但实地测试成本高,现有仿真对罕见事故的保真度不足。事故报告提供了真实交通事故的丰富细节,是大语言模型与视觉语言模型生成高保真场景的宝贵资源。然而,现有模型常因上下文抑制而偏离实际事故特征。为此,本文提出SG-CADVLM框架,结合上下文感知解码与多模态输入处理,从事故报告中生成安全关键场景。该框架在生成道路几何与车辆轨迹的同时,有效抑制视觉语言模型的幻觉。实验表明,相较于基线方法31.2%的生成成功率,SG-CADVLM实现88.1%的危急与高风险场景生成率,提升182%,并可生成可用于自动驾驶测试的可执行仿真。
原文摘要 · Abstract (English)
Autonomous Vehicle (AV) requires rigorous testing in safety-critical scenarios for safety validation, yet its validation is hindered by the high cost of field testing and the lack of fidelity in current simulations for rare safety-critical events. Crash reports offer rich and authentic specifications of real-world accident dynamics, making them a promising resource for Large Language Models and Vision-Language models to generate high-fidelity scenarios. However, the existing models frequently deviate from actual accident characteristics due to context suppression. To address these limitations, this paper presents SG-CADVLM, a framework integrateing Context-Aware Decoding with multimodal input processing to generate safety-critical scenarios from crash reports. The framework mitigates the hallucination of VLMs while generating road geometry and vehicle trajectories simultaneously. The experimental results demonstrate that SG-CADVLM generates combined critical and high-risk scenarios at a rate of 88.1% compared to 31.2% for the baseline methods, representing a 182% improvement, while producing executable simulations for autonomous vehicle testing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。