用语言引导3D轨迹规划,让无人车更懂复杂地形。
Reasoning About Traversability: Language-Guided Off-Road 3D Trajectory Planning

- 将语言标注对齐车辆动作与地形,提升语义理解精度
- 轨迹误差降为0.97米,地形合规率提升至0.644
- 适合自动驾驶在非结构化地形中做智能路径规划
视觉语言模型(VLM)虽能实现端到端自主驾驶的高层语义推理,尤其在非结构化环境表现突出,但现有越野数据集的语言标注与车辆行为及地形几何存在弱对齐问题。为此,我们提出一种语言精炼框架,将标注重构为动作对齐的样本对,使VLM可从单张图像直接生成优化后的场景描述与3D未来轨迹。为进一步强化地形感知,引入偏好优化策略,构建几何相关的难例负样本,并显式惩罚与局部高程剖面不一致的轨迹。此外,提出针对越野场景的评估指标,量化可通行性合规度与高程一致性,弥补传统道路评价的不足。在ORAD-3D基准测试中,方法将平均轨迹误差从1.01米降至0.97米,可通行性合规度由0.621提升至0.644,高程不一致性从0.428降至0.322,验证了动作对齐监督与地形感知优化的有效性。
原文摘要 · Abstract (English)
While Vision-Language Models (VLMs) enable high-level semantic reasoning for end-to-end autonomous driving, particularly in unstructured environments, existing off-road datasets suffer from language annotations that are weakly aligned with vehicle actions and terrain geometry. To address this misalignment, we propose a language refinement framework that restructures annotations into action-aligned pairs, enabling a VLM to generate refined scene descriptions and 3D future trajectories directly from a single image. To further encourage terrain-aware planning, we introduce a preference optimization strategy that constructs geometry-aware hard negatives and explicitly penalizes trajectories inconsistent with local elevation profiles. Furthermore, we propose off-road-specific metrics to quantify traversability compliance and elevation consistency, addressing the limitations of conventional on-road evaluation. Experiments on the ORAD-3D benchmark demonstrate that our approach reduces average trajectory error from 1.01m to 0.97m, improves traversability compliance from 0.621 to 0.644, and decreases elevation inconsistency from 0.428 to 0.322, highlighting the efficacy of action-aligned supervision and terrain-aware optimization for robust off-road driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。