推理模型常言行不一,真实驾驶行为与推理不符。
Reasoning models do not yet follow their reasoning in autonomous driving: The KITScenes LongTail Dataset
- 构建罕见场景数据集,量化推理与执行的一致性
- 多数模型推理正确但行动偏离,导致信任危机
- 让模型按推理执行能显著提升规划效果
处理罕见事件是自动驾驶的核心挑战。推理模型通过生成显式推理链再执行动作,有望应对此类情况。但我们发现,这些模型经常不遵循自身推理:其推理中声明的动作与实际执行动作存在显著偏差。为此,我们构建了KITScenes LongTail数据集,专门用于量化推理与动作之间的语义一致性。结果显示,当前通用及领域特定模型普遍存在不一致现象。令人惊讶的是,当推理与执行发生分歧时,推理内容通常更准确:若从推理链提取动作并由运动学模型执行,可恢复一致性并显著改善运动规划性能。这表明当前模型的推理能力已超越其行为表现,而让行为忠实于推理是实现可信自动驾驶的前提。数据集与评估平台已公开:https://hf.co/datasets/kit-mrt/kitscenes-longtail
原文摘要 · Abstract (English)
Handling rare events is the central open challenge in autonomous driving. Reasoning models, which generate explicit chains of reasoning before acting, promise to generalize to such events. Here we show that these models frequently do not follow their own reasoning: the actions they state in their reasoning often diverge from the actions they ultimately execute. We introduce KITScenes LongTail, a curated dataset of rare driving scenarios to quantify this divergence through a measure of semantic reasoning-action coherence. We find that incoherence is widespread across current general-purpose and domain-specific models. Strikingly, when reasoned actions and executed actions disagree, the reasoning is usually right: extracting actions from the reasoning trace and executing them through a kinematic model enforces coherence and typically improves motion planning. These results suggest that the reasoning capabilities of current models exceed what their actions reveal, and that coherently acting on stated reasoning is a prerequisite for trustworthy autonomous driving. Our dataset and evaluation space are available at: https://hf.co/datasets/kit-mrt/kitscenes-longtail
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。