arXiv:2505.17685cs.CV2025-05NeurIPS被引 215

让自动驾驶模型'视觉思考',用时空图像推理未来路况。

FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving

  • 用生成的未来场景作为视觉时空思维链,融合空间结构与时间演化。
  • 在nuScenes和NAVSIM上轨迹更准、碰撞更少,视频生成FID表现优。
  • 适合关注自动驾驶规划与视觉推理融合的研究者或工程师。

视觉-语言-动作(VLA)模型在端到端自动驾驶中潜力巨大,但其推理常受限于文本型思维链(CoT),导致视觉信息符号化压缩,造成感知与规划间的模态鸿沟,模糊时空关系并丢失细粒度线索。本文提出FSDrive框架,使VLA能够通过新型视觉时空思维链实现‘视觉思考’。该框架首先作为世界模型,生成包含预测背景与未来车道线、3D物体框等物理合理先验的统一未来帧;这一想象中的场景即为视觉时空思维链,单一体现空间结构与时间演进。同一VLA随后作为逆动力学模型,基于当前观测与该视觉思维链规划轨迹。通过统一预训练范式扩展模型视觉词汇表,并联合优化语义理解(VQA)与未来帧预测。采用渐进式课程学习,先生成结构先验以符合物理规律,再渲染完整场景。在nuScenes和NAVSIM上的评估显示,FSDrive显著提升轨迹准确率并减少碰撞,同时以轻量级自回归模型实现具有竞争力的视频生成FID表现,并在DriveLM上推进场景理解能力。结果证实,视觉时空思维链有效弥合感知-规划鸿沟,推动更安全、更具前瞻性的自动驾驶。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models offer significant potential for end-to-end driving, yet their reasoning is often constrained by textual Chains-of-Thought (CoT). This symbolic compression of visual information creates a modality gap between perception and planning by blurring spatio-temporal relations and discarding fine-grained cues. We introduce FSDrive, a framework that empowers VLAs to "think visually" using a novel visual spatio-temporal CoT. FSDrive first operates as a world model, generating a unified future frame that combines a predicted background with explicit, physically-plausible priors like future lane dividers and 3D object boxes. This imagined scene serves as the visual spatio-temporal CoT, capturing both spatial structure and temporal evolution in a single representation. The same VLA then functions as an inverse-dynamics model to plan trajectories conditioned on current observations and this visual CoT. We enable this with a unified pre-training paradigm that expands the model's vocabulary with visual tokens and jointly optimizes for semantic understanding (VQA) and future-frame prediction. A progressive curriculum first generates structural priors to enforce physical laws before rendering the full scene. Evaluations on nuScenes and NAVSIM show FSDrive improves trajectory accuracy and reduces collisions, while also achieving competitive FID for video generation with a lightweight autoregressive model and advancing scene understanding on DriveLM. These results confirm that our visual spatio-temporal CoT bridges the perception-planning gap, enabling safer, more anticipatory autonomous driving. Code is available at https://github.com/MIV-XJTU/FSDrive.

自动驾驶视觉推理时空建模VLA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。