arXiv:2602.17677cs.LGcs.CL2026-02被引 1

解决自动驾驶视觉语言模型问答中的文本偏见问题

Reducing Text Bias in Synthetically Generated MCQAs for VLMs in Autonomous Driving

  • 分离正确答案与语言特征,避免模型依赖文字套路
  • 盲测准确率从高出随机66.9%降至仅高2.9%
  • 适合评估真实视觉理解能力的VLM基准构建

多选题问答(MCQA)基准是衡量自动驾驶场景下视觉语言模型(VLM)性能的公认标准。然而我们发现,合成生成的MCQA极易受隐藏文本线索影响,使模型可利用语言模式而非视觉上下文作答。实验显示,仅在合成数据上微调的VLM,在无视觉输入时仍能达到接近人工标注基准的准确率。本文提出的方法将盲测准确率从高于随机66.9%降低至仅高2.9%,基本消除可被利用的语言捷径。通过解耦正确答案与语言特征,并采用课程学习策略,强制模型依赖视觉基础,确保性能真实反映感知理解能力。

原文摘要 · Abstract (English)

Multiple Choice Question Answering (MCQA) benchmarks are an established standard for measuring Vision Language Model (VLM) performance in driving tasks. However, we observe the known phenomenon that synthetically generated MCQAs are highly susceptible to hidden textual cues that allow models to exploit linguistic patterns rather than visual context. Our results show that a VLM fine-tuned on such data can achieve accuracy comparable to human-validated benchmarks even without visual input. Our proposed method reduces blind accuracy from +66.9% above random to +2.9%, eliminating the vast majority of exploitable textual shortcuts. By decoupling the correct answer from linguistic artifacts and employing a curriculum learning strategy, we force the model to rely on visual grounding, ensuring that performance accurately reflects perceptual understanding.

视觉语言模型自动驾驶文本偏见测评基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。