arXiv:2602.10771cs.CVcs.RO2026-02

测试视觉语言模型对骑行者视角的交通理解能力,发现现有模型仍有明显短板。

From Steering to Pedalling: Do Autonomous Driving VLMs Generalize to Cyclist-Assistive Spatial Perception and Planning?

  • 构建骑行者视角的诊断基准CyclingVQA,评估感知与决策能力。
  • 31+模型中多数表现尚可,但对骑行专属交通信号理解不足。
  • 驾驶专用模型反而不如通用模型,提示任务迁移存在局限。

骑行者在城市交通中常面临安全危机,亟需辅助系统支持安全决策。近期视觉语言模型(VLMs)在自动驾驶基准上表现优异,暗示其具备交通理解与导航推理潜力。然而,现有评估多以车辆为中心,缺乏从骑行者视角出发的评测。为此,我们提出CyclingVQA,一个针对骑行者视角的诊断基准,用于检验感知、时空理解及交通规则到车道的关联推理能力。我们评估了31+种主流VLMs,涵盖通用型、空间增强型及自动驾驶专用模型,发现当前模型虽具潜力,但在解读骑行者专属交通线索以及将标识与正确车道关联方面仍存明显不足。值得注意的是,部分驾驶专用模型的表现甚至低于强通用模型,表明从车辆中心训练向骑行辅助场景的迁移效果有限。通过系统性错误分析,我们识别出若干重复出现的失效模式,为开发更有效的骑行辅助智能系统提供指引。

原文摘要 · Abstract (English)

Cyclists often encounter safety-critical situations in urban traffic, highlighting the need for assistive systems that support safe and informed decision-making. Recently, vision-language models (VLMs) have demonstrated strong performance on autonomous driving benchmarks, suggesting their potential for general traffic understanding and navigation-related reasoning. However, existing evaluations are predominantly vehicle-centric and fail to assess perception and reasoning from a cyclist-centric viewpoint. To address this gap, we introduce CyclingVQA, a diagnostic benchmark designed to probe perception, spatio-temporal understanding, and traffic-rule-to-lane reasoning from a cyclist's perspective. Evaluating 31+ recent VLMs spanning general-purpose, spatially enhanced, and autonomous-driving-specialized models, we find that current models demonstrate encouraging capabilities, while also revealing clear areas for improvement in cyclist-centric perception and reasoning, particularly in interpreting cyclist-specific traffic cues and associating signs with the correct navigational lanes. Notably, several driving-specialized models underperform strong generalist VLMs, indicating limited transfer from vehicle-centric training to cyclist-assistive scenarios. Finally, through systematic error analysis, we identify recurring failure modes to guide the development of more effective cyclist-assistive intelligent systems.

视觉语言模型骑行辅助交通理解多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。