用自然语言描述路况,精准识别驾驶风险并给出安全建议。
DriveSafe: A Framework for Risk Detection and Safety Suggestions in Driving Scenarios

- 生成带空间信息的多模态场景描述,支撑精细风险判断。
- 在DRAMA数据集上超越零样本模型与已有方法,显著提升准确率。
- 适合自动驾驶系统开发与交通安全隐患分析人员使用。
自主车辆在安全关键环境中运行需具备全面的情境感知能力,以识别并缓解潜在风险。尽管近期多模态大语言模型(MLLMs)在通用视觉-语言任务中表现良好,但我们的研究发现,零样本的MLLM在细粒度、空间定位的风险评估任务中仍逊于领域专用方法。为此,我们提出DriveSafe框架,通过结构化自然语言描述实现风险感知的场景理解。该方法首先生成融合运动、空间和深度线索的具空间定位的图像描述;随后基于这些描述进行下游风险评估,明确识别危险物体及其位置,并推断其隐含的不安全行为,进而提供可操作的安全建议。为进一步提升性能,我们采用描述-风险配对数据微调轻量级适配模块,高效注入领域知识至基础大模型。通过基于显式语言表征的条件化风险评估,DriveSafe在多个方面显著优于零样本MLLM及先前领域专用基线。在DRAMA基准上的充分实验验证了其最先进的性能,消融实验也证实了核心设计的有效性。
原文摘要 · Abstract (English)
Comprehensive situational awareness is essential for autonomous vehicles operating in safety-critical environments, as it enables the identification and mitigation of potential risks. Although recent Multimodal Large Language Models (MLLMs) have shown promise on general vision-language tasks, our findings indicate that zero-shot MLLMs still underperform compared to domain-specific methods in fine-grained, spatially grounded risk assessment. To address this gap, we propose DriveSafe, a framework for risk-aware scene understanding that leverages structured natural language descriptions. Specifically, our method first generates spatially grounded captions enriched with multimodal context, including motion, spatial, and depth cues. These captions are then used for downstream risk assessment, explicitly identifying hazardous objects, their locations, and the unsafe behaviors they imply, followed by actionable safety suggestions. To further improve performance, we employ caption-risk pairings to fine-tune a lightweight adapter module, efficiently injecting domain-specific knowledge into the base LLM. By conditioning risk assessment on explicit language-based scene representations, DriveSafe achieves significant gains over both zero-shot MLLMs and prior domain-specific baselines. Exhaustive experiments on the DRAMA benchmark demonstrate state-of-the-art performance, while ablation studies validate the effectiveness of our key design choices. Project page: https://cvit.iiit.ac.in/ research/projects/cvit-projects/drivesafe
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。