自动驾驶中的推理需与真实动作同步,而非仅靠文本思考。
Beyond Textual Chain-of-Thought: A Survey on Action-Grounded Reasoning in Autonomous Driving

- 以中间表示形式为轴,分类130篇方法论文的推理模式。
- 发现真实世界可验证的动态表征是智能驾驶推理的关键。
- 适合研究自动驾驶决策与安全验证的学者参考。
思维链(Chain-of-thought, CoT)推理通过生成中间步骤来驱动生成模型,但在自动驾驶中,输出是连续动作,因此其推理必须与物理世界的时空结构一致。本文系统考察从文本思维链到动作锚定推理的范式转变。综述171篇文献,包括130篇方法论文及41篇基准、数据集、综述与分析论文,提出以表示为中心的分类体系,将130种方法归为四类:语言型、视觉空间型、隐式动态型和外部化推理,进一步细分为13个子类型,对应不同关注区域。研究表明,自动驾驶智能体推理的前沿在于可真实锚定、实时关联动作、并能在高安全要求系统中验证的中间表征。项目页面:https://github.com/tangzhengxu/awesome-av-cot。
原文摘要 · Abstract (English)
Chain-of-thought (CoT) reasoning powers generative models by eliciting intermediate steps before producing an answer. In autonomous driving, the answer is a continuous action. Thus its reasoning must share the same spatiotemporal structure as the physical world. This survey studies the resulting shift from textual CoT to action-grounded reasoning. Surveying 171 papers, including 130 method papers and 41 benchmarks, datasets, surveys, and analysis papers, we propose a representation-centered taxonomy that treats the form of the intermediate state as the organizing axis. We systematize the 130 methods into four categories: language-based, visual-spatial, latent-dynamic, and externalized reasoning, further divided into 13 subtypes tied to distinct regions of interests. Our synthesis shows that the open frontier of reasoning in driving agents lies in intermediate representations that can be grounded in the real world, coupled to real-time action, and verified under safety-critical systems. Project page: https://github.com/tangzhengxu/awesome-av-cot.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。