从规则驾驶转向端到端大模型,自动驾驶正迈向全自主时代。
The Era of End-to-End Autonomy: Transitioning from Rule-Based Driving to Large Driving Models
- 用大模型直接从传感器输入生成驾驶动作,跳过传统分模块流程。
- 2026年起多家厂商将部署需人监督的端到端系统,可处理复杂路况。
- 适合关注自动驾驶商业化落地与具身智能发展的从业者。
自动驾驶正从传统的分模块感知-规划-控制架构,转向端到端(E2E)学习系统。本文梳理了从经典架构到大型驾驶模型(LDMs)的发展历程,这些模型能将原始传感器输入直接映射为驾驶动作。研究聚焦特斯拉FSD V12/V14、Rivian统一智驾平台、NVIDIA Cosmos及新兴商业化无人出租车部署,分析其架构设计、部署策略、安全考量与产业影响。一个关键趋势是监督式端到端驾驶(常称FSD(监督)或L2++),多家厂商计划自2026年起部署,可在复杂环境中完成大部分动态驾驶任务(DDT),但需人类持续监督,司机角色转为安全监护。早期运行数据表明,端到端学习更有效应对真实世界长尾场景,已成为主流商业策略。此外,此类架构进步也可能扩展至其他具身智能系统,如人形机器人。
原文摘要 · Abstract (English)
Autonomous driving is undergoing a shift from modular rule based pipelines toward end to end (E2E) learning systems. This paper examines this transition by tracing the evolution from classical sense perceive plan control architectures to large driving models (LDMs) capable of mapping raw sensor input directly to driving actions. We analyze recent developments including Tesla's Full Self Driving (FSD) V12 V14, Rivian's Unified Intelligence platform, NVIDIA Cosmos, and emerging commercial robotaxi deployments, focusing on architectural design, deployment strategies, safety considerations and industry implications. A key emerging product category is supervised E2E driving, often referred to as FSD (Supervised) or L2 plus plus, which several manufacturers plan to deploy from 2026 onwards. These systems can perform most of the Dynamic Driving Task (DDT) in complex environments while requiring human supervision, shifting the driver's role to safety oversight. Early operational evidence suggests E2E learning handles the long tail distribution of real world driving scenarios and is becoming a dominant commercial strategy. We also discuss how similar architectural advances may extend beyond autonomous vehicles (AV) to other embodied AI systems, including humanoid robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。