用语言模型统一感知、决策,让自动驾驶更懂路况、更易解释。
A Unified Perception-Language-Action Framework for Adaptive Autonomous Driving
- 融合多传感器与大语言模型,实现感知到决策的端到端理解。
- 在施工路口场景中轨迹追踪误差降低32%,规划更适应复杂变化。
- 适合追求安全可解释性的自动驾驶研发者,尤其关注长尾场景应对。
自动驾驶系统在复杂开放环境中的自适应性、鲁棒性和可解释性仍面临挑战,根源在于架构碎片化、对新场景泛化能力弱以及感知中语义提取不足。为此,我们提出统一的感知-语言-行动(PLA)框架,融合摄像头、激光雷达和雷达的多传感器信息,结合大语言模型(LLM)增强的视觉-语言-行动(VLA)架构,采用GPT-4.1作为推理核心。该框架将低层感知与高层语境推理紧密结合,通过自然语言实现语义理解与决策联动,支持上下文感知、可解释且安全有界的自动驾驶。在城市交叉口含施工区的场景测试中,轨迹跟踪、速度预测与自适应规划性能均显著优于基准方法。结果表明,语言增强的认知框架有望提升自动驾驶系统的安全性、可解释性与可扩展性。
原文摘要 · Abstract (English)
Autonomous driving systems face significant challenges in achieving human-like adaptability, robustness, and interpretability in complex, open-world environments. These challenges stem from fragmented architectures, limited generalization to novel scenarios, and insufficient semantic extraction from perception. To address these limitations, we propose a unified Perception-Language-Action (PLA) framework that integrates multi-sensor fusion (cameras, LiDAR, radar) with a large language model (LLM)-augmented Vision-Language-Action (VLA) architecture, specifically a GPT-4.1-powered reasoning core. This framework unifies low-level sensory processing with high-level contextual reasoning, tightly coupling perception with natural language-based semantic understanding and decision-making to enable context-aware, explainable, and safety-bounded autonomous driving. Evaluations on an urban intersection scenario with a construction zone demonstrate superior performance in trajectory tracking, speed prediction, and adaptive planning. The results highlight the potential of language-augmented cognitive frameworks for advancing the safety, interpretability, and scalability of autonomous driving systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。