arXiv:2605.10177cs.CVcs.AI2026-05

用多模态Transformer构建3D可行驶性表征,让自动驾驶决策更稳定高效

MTA-RL: Robust Urban Driving via Multi-modal Transformer-based 3D Affordances and Reinforcement Learning

论文配图:MTA-RL: Robust Urban Driving via Multi-modal Transformer-based 3D Affordances and Reinforcement Learning
图 1 · 摘自论文原文
  • 通过融合图像与点云的Transformer生成几何感知的驾驶语义表征
  • 在复杂城市场景中实现9.0%的路线完成率提升和83.7%的违规距离改善
  • 适合追求高鲁棒性与零样本泛化的自动驾驶系统研究者

可靠的城市自动驾驶需要在密集交互下具备稳定的三维场景理解与决策能力。现有端到端模型缺乏可解释性,而模块化流程则存在脆弱接口导致的误差传播问题。本文提出MTA-RL,首个通过多模态Transformer驱动的3D可行驶性与强化学习融合框架。不同于直接回归动作的传统融合模型,该方法利用Transformer融合RGB图像与LiDAR点云,预测显式的、几何感知的可行驶性表征。这些结构化表示作为紧凑的观测空间,使强化学习策略仅基于预测的驾驶语义进行操作,显著提升采样效率与稳定性。在CARLA Town01-03中不同车流密度(20-60辆背景车辆)下的大量测试表明,MTA-RL持续优于现有先进基线。仅在Town03训练的模型,在未见过的城镇中展现出优越的零样本泛化能力:路线完成率提升9.0%,总行驶距离增加11.0%,单位违规距离提升83.7%。消融实验进一步验证了多模态融合与奖励设计的关键作用,显著优于仅使用图像或未设计奖励的变体,证明其在鲁棒城市自动驾驶中的有效性。

原文摘要 · Abstract (English)

Robust urban autonomous driving requires reliable 3D scene understanding and stable decision-making under dense interactions. However, existing end-to-end models lack interpretability, while modular pipelines suffer from error propagation across brittle interfaces. This paper proposes MTA-RL, the first framework that bridges perception and control through Multi-modal Transformer-based 3D Affordances and Reinforcement Learning (RL). Unlike previous fusion models that directly regress actions, RGB images and LiDAR point clouds are fused using a transformer architecture to predict explicit, geometry-aware affordance representations. These structured representations serve as a compact observation space, enabling the RL policy to operate purely on predicted driving semantics, which significantly improves sample efficiency and stability. Extensive evaluations in CARLA Town01-03 across varying densities (20-60 background vehicles) show that MTA-RL consistently outperforms state-of-the-art baselines. Trained solely on Town03, our method demonstrates superior zero-shot generalization in unseen towns, achieving up to a 9.0% increase in Route Completion, an 11.0% increase in Total Distance, and an 83.7% improvement in Distance Per Violation. Furthermore, ablation studies confirm that our multi-modal fusion and reward shaping are critical, significantly outperforming image-only and unshaped variants, demonstrating the effectiveness of MTA-RL for robust urban autonomous driving.

自动驾驶多模态融合强化学习3D感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。