提出可评估视觉语言模型驾驶决策是否真懂人类理由的框架
CARE Drive A Framework for Evaluating Reason-Responsiveness of Vision Language Models in Automated Driving
- 通过控制上下文变化对比模型决策,判断人类理由是否真实影响行为
- 在骑车人超车场景中,显式理由显著提升决策与专家建议的一致性
- 适合关注自动驾驶模型可解释性与安全可信性的研究者
基础模型(包括视觉语言模型)在自动驾驶中被用于场景理解、行为推荐和生成自然语言解释。然而,现有评估方法主要关注结果性能(如安全性与轨迹准确性),未检验模型决策是否反映人类相关考量。这导致无法判断其解释是真实因果推理还是事后合理化,在安全关键领域可能造成虚假信心。为此,我们提出 CARE Drive:一种面向自动驾驶的上下文感知理由评估框架,用于评估视觉语言模型的因果响应能力。该框架采用两阶段流程:首先进行提示校准以确保输出稳定;再通过系统性上下文扰动,测量模型对安全余量、社会压力、效率约束等人类理由的敏感度。在涉及规范冲突的骑车人超车场景中,结果显示显式人类理由显著影响模型决策,提升与专家推荐行为的一致性。但响应程度因上下文因素而异,表明模型对不同理由的敏感性不均。这些发现证明,无需修改模型参数即可系统评估基础模型的理由响应能力。
原文摘要 · Abstract (English)
Foundation models, including vision language models, are increasingly used in automated driving to interpret scenes, recommend actions, and generate natural language explanations. However, existing evaluation methods primarily assess outcome based performance, such as safety and trajectory accuracy, without determining whether model decisions reflect human relevant considerations. As a result, it remains unclear whether explanations produced by such models correspond to genuine reason responsive decision making or merely post hoc rationalizations. This limitation is especially significant in safety critical domains because it can create false confidence. To address this gap, we propose CARE Drive, Context Aware Reasons Evaluation for Driving, a model agnostic framework for evaluating reason responsiveness in vision language models applied to automated driving. CARE Drive compares baseline and reason augmented model decisions under controlled contextual variation to assess whether human reasons causally influence decision behavior. The framework employs a two stage evaluation process. Prompt calibration ensures stable outputs. Systematic contextual perturbation then measures decision sensitivity to human reasons such as safety margins, social pressure, and efficiency constraints. We demonstrate CARE Drive in a cyclist overtaking scenario involving competing normative considerations. Results show that explicit human reasons significantly influence model decisions, improving alignment with expert recommended behavior. However, responsiveness varies across contextual factors, indicating uneven sensitivity to different types of reasons. These findings provide empirical evidence that reason responsiveness in foundation models can be systematically evaluated without modifying model parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。