用推理链训练小模型,让自动驾驶问答更可靠。
ReasonDrive: Efficient Visual Question Answering for Autonomous Vehicles with Reasoning-Enhanced Small Vision-Language Models
- 在微调中引入结构化推理链,提升模型决策透明度。
- 小模型(如Llama3.2-11B-reason)在DriveLM上准确率显著提升。
- 适合关注自动驾驶可解释性与安全性的研究者。
视觉语言模型(VLMs)在自动驾驶中前景广阔,但缺乏透明的推理能力,影响安全性。本文研究在微调阶段显式建模推理是否能提升驾驶决策任务表现。利用GPT-4o,我们基于类别特定提示策略,为DriveLM基准中的驾驶场景生成结构化推理链。在多个小型VLM家族(Llama 3.2、Llava 1.5、Qwen 2.5VL)上对比了基于推理的微调、仅答案微调和基线指令微调模型。结果表明,基于推理的微调持续优于其他方法,其中Llama3.2-11B-reason表现最佳。使用推理微调的模型在准确率和文本生成质量上均有显著提升,说明显式推理增强了模型对驾驶决策的内部表征。这些发现凸显了安全关键领域中透明决策过程的重要性,并为构建更具可解释性的自动驾驶系统提供了新方向。
原文摘要 · Abstract (English)
Vision-language models (VLMs) show promise for autonomous driving but often lack transparent reasoning capabilities that are critical for safety. We investigate whether explicitly modeling reasoning during fine-tuning enhances VLM performance on driving decision tasks. Using GPT-4o, we generate structured reasoning chains for driving scenarios from the DriveLM benchmark with category-specific prompting strategies. We compare reasoning-based fine-tuning, answer-only fine-tuning, and baseline instruction-tuned models across multiple small VLM families (Llama 3.2, Llava 1.5, and Qwen 2.5VL). Our results demonstrate that reasoning-based fine-tuning consistently outperforms alternatives, with Llama3.2-11B-reason achieving the highest performance. Models fine-tuned with reasoning show substantial improvements in accuracy and text generation quality, suggesting explicit reasoning enhances internal representations for driving decisions. These findings highlight the importance of transparent decision processes in safety-critical domains and offer a promising direction for developing more interpretable autonomous driving systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。