通过侧车监督提升自动驾驶模型对周边车辆的感知能力
Mind the Privileged-to-Camera Gap: Actor-Centric Sidecar Supervision for Camera-First Open-Loop Waypoint Prediction

- 引入侧车标签监督车辆级表征,弥补仅靠路径点监督的不足
- 路径预测误差降低32.6%,尤其在有弱势道路使用者时效果显著
- 适合关注自动驾驶感知与规划融合的开发者和研究者
以摄像头为主的自动驾驶模型从图像、自身状态和路线指令中预测未来自身路径点,但仅依赖路径点监督无法显式指导附近交通参与者(道路使用者)的表征学习。本文将此问题视为开环路径预测中的监督表征学习。部署模型在推理时使用多视角RGB图像、自身状态和路线指令;训练阶段利用模拟器生成的侧车标签,监督车辆定位、相对于历史轨迹的隐式相关性以及选定车辆的短时运动。这些标签不作为推理输入。在路线分离的评估设置下,采用相同架构、优化器、验证指标、检查点选择和三组随机种子。纯路径点监督的基线模型最终位移误差为1.815±0.02米,无教师的非侧车控制组为1.716±0.02米。加入道路使用者侧车监督(RU-sidecar)后,最终位移误差降至1.223±0.01米,相比基线降低32.6%,相比对照组降低28.7%。在1445/1494条路线中优于基线,在1417/1494条路线中优于对照组。分片分析显示所有非空子集均有提升,含至少四名有效侧车车辆样本下降29.1%,存在弱势道路使用者时下降30.0%。可选的模拟器状态教师对齐进一步降至1.186±0.15米,但种子间波动较大,次之。非部署型模拟器诊断仍更强,表明存在特权信息到摄像头的差距。证据限于开环仿真诊断。
原文摘要 · Abstract (English)
Camera-first autonomous-driving models predict future ego waypoints from images, ego-state features, and route commands, but waypoint supervision alone does not explicitly supervise actor-level representations of nearby road users. We study this as supervised representation learning for open-loop waypoint prediction. The deployable model uses multi-view RGB, ego state, and route command at inference. During training, simulator-derived sidecar labels supervise actor grounding, privileged hindsight actor relevance relative to the logged ego trajectory, and selected-actor short-horizon motion; these labels are never inference inputs. We evaluate route-disjoint splits with matched architecture, optimizer, validation criterion, checkpoint selection, and three seeds. A plain waypoint-only RGB baseline obtains 1.815$\pm$0.02 m final displacement error (FDE), and the matched no-teacher non-sidecar RGB control obtains 1.716$\pm$0.02 m. Road-user sidecar supervision (RU-sidecar) reduces FDE to 1.223$\pm$0.01 m, a 32.6% reduction over the plain baseline and 28.7% over the matched no-teacher non-sidecar RGB control. It improves over the plain baseline on 1445/1494 routes and over the matched no-teacher non-sidecar RGB control on 1417/1494 routes. Actor-conditioned slices show gains in all nonempty subsets, including 29.1% reduction for samples with at least four valid sidecar actors and 30.0% when a vulnerable road user is present. Optional simulator-state teacher alignment reaches 1.186$\pm$0.15 m FDE, but higher seed variability makes it secondary. Non-deployable simulator-state diagnostics remain stronger, indicating a privileged-to-camera gap. The evidence is limited to open-loop simulation diagnostics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。