arXiv:2601.03519cs.RO2026-01被引 2

用视觉提示+思维链提升越野自动驾驶规划精度

A Vision-Language-Action Model with Visual Prompt for OFF-Road Autonomous Driving

  • 设计视觉提示块,用语义掩码增强模型空间理解
  • 引入自洽思维链,使轨迹规划错误率降为6.56%
  • 适合研究越野自动驾驶与多模态决策的开发者

在非铺装地形中高效进行路径规划对自动驾驶车辆构成重大挑战,通常需复杂的多步骤流程。传统方法在动态环境中适应性差。为此,本文提出OFF-EMMA,一种新型端到端多模态框架,旨在解决视觉-语言-动作(VLA)模型在非铺装场景下空间感知不足和推理不稳的问题。该框架通过设计视觉提示模块显式标注输入图像,并引入带有自洽性的思维链(COT-SC)推理策略,提升路径规划的准确性和鲁棒性。视觉提示模块利用语义分割掩码作为视觉提示,增强预训练视觉-语言模型对复杂地形的空间理解能力。COT-SC策略通过多路径推理机制有效缓解异常值对规划性能的影响。在RELLIS-3D非铺装数据集上的实验表明,OFF-EMMA显著优于现有方法,将基于Qwen骨干模型的平均L2误差降低13.3%,并将失败率从16.52%降至6.56%。

原文摘要 · Abstract (English)

Efficient trajectory planning in off-road terrains presents a formidable challenge for autonomous vehicles, often necessitating complex multi-step pipelines. However, traditional approaches exhibit limited adaptability in dynamic environments. To address these limitations, this paper proposes OFF-EMMA, a novel end-to-end multimodal framework designed to overcome the deficiencies of insufficient spatial perception and unstable reasoning in visual-language-action (VLA) models for off-road autonomous driving scenarios. The framework explicitly annotates input images through the design of a visual prompt block and introduces a chain-of-thought with self-consistency (COT-SC) reasoning strategy to enhance the accuracy and robustness of trajectory planning. The visual prompt block utilizes semantic segmentation masks as visual prompts, enhancing the spatial understanding ability of pre-trained visual-language models for complex terrains. The COT- SC strategy effectively mitigates the error impact of outliers on planning performance through a multi-path reasoning mechanism. Experimental results on the RELLIS-3D off-road dataset demonstrate that OFF-EMMA significantly outperforms existing methods, reducing the average L2 error of the Qwen backbone model by 13.3% and decreasing the failure rate from 16.52% to 6.56%.

自动驾驶视觉提示多模态路径规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。