用3D基础模型先验提升自动驾驶在视角变化下的稳定性
Towards Viewpoint-Robust End-to-End Autonomous Driving with 3D Foundation Model Priors
- 用3D位置信息作为像素级位置嵌入,融合几何特征
- 在视角扰动测试中,俯仰和高度变化下性能下降减少
- 适合需要强视角鲁棒性的自动驾驶系统开发
在摄像头视角变化下实现稳健的轨迹规划对可扩展的端到端自动驾驶至关重要。现有模型往往严重依赖训练时见过的摄像头视角。本文提出一种无需数据增强的方法,利用3D基础模型提供的几何先验。该方法将从深度估计得到的每像素3D位置作为位置嵌入注入,并通过交叉注意力融合中间几何特征。在VR-Drive摄像头视角扰动基准上的实验表明,多数扰动条件下性能下降减少,俯仰和高度扰动下改进明显。纵向平移下的增益较小,表明还需更视角无关的融合机制以增强对摄像头视角变化的鲁棒性。
原文摘要 · Abstract (English)
Robust trajectory planning under camera viewpoint changes is important for scalable end-to-end autonomous driving. However, existing models often depend heavily on the camera viewpoints seen during training. We investigate an augmentation-free approach that leverages geometric priors from a 3D foundation model. The method injects per-pixel 3D positions derived from depth estimates as positional embeddings and fuses intermediate geometric features through cross-attention. Experiments on the VR-Drive camera viewpoint perturbation benchmark show reduced performance degradation under most perturbation conditions, with clear improvements under pitch and height perturbations. Gains under longitudinal translation are smaller, suggesting that more viewpoint-agnostic integration is needed for robustness to camera viewpoint changes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。