为高铁自动驾驶设计首个视觉认知评测基准与高效可解释推理框架
RailVQA: A Benchmark and Framework for Efficient Interpretable Visual Cognition in Automatic Train Operation
- 构建小模型与大模型协作的透明三模块架构,结合自适应采样提升效率
- 在2万张单帧和1168段视频上验证,显著提升跨场景泛化与可解释性
- 适合自动驾驶、铁路安全等高可靠场景的研究者与工程团队使用
随着自动列车运行(ATO)向GoA4及以上级别演进,其对高效可靠的司机室视角视觉感知与面向决策的推理能力需求日益增强,以确保在复杂动态铁路环境中的安全运行。然而现有方法多聚焦基础感知,对罕见但关键的安全边缘案例泛化能力差,且缺乏高层推理与规划能力。尽管近期大型多模态模型(LMMs)展现出强大泛化与认知能力,但在安全关键型ATO中受限于高计算成本与幻觉风险。同时,系统评估认知能力的领域专用基准仍属空白。为此,我们提出RailVQA-bench,首个针对ATO司机室视觉认知的VQA基准,包含20,000个单帧与1,168个视频问答对,用于评估静态与动态场景下的认知泛化与可解释性。此外,我们提出RailVQA-CoM,一种通过透明三模块架构与自适应时间采样实现大-小模型协同的框架,结合小模型效率与大模型认知能力,显著提升感知泛化、推理效率与规划能力。实验表明,该方法在性能、可解释性、效率及跨域泛化方面均有显著提升。代码与数据集将公开于https://cybereye-bjtu.github.io/RailVQA.html。
原文摘要 · Abstract (English)
As Automatic Train Operation (ATO) advances toward GoA4 and beyond, it increasingly depends on efficient, reliable cab-view visual perception and decision-oriented inference to ensure safe operation in complex and dynamic railway environments. However, existing approaches focus primarily on basic perception and often generalize poorly to rare yet safety-critical corner cases. They also lack the high-level reasoning and planning capabilities required for operational decision-making. Although recent Large Multi-modal Models (LMMs) show strong generalization and cognitive capabilities, their use in safety-critical ATO is hindered by high computational cost and hallucination risk. Meanwhile, reliable domain-specific benchmarks for systematically evaluating cognitive capabilities are still lacking. To address these gaps, we introduce RailVQA-bench, the first VQA benchmark for cab-view visual cognition in ATO, comprising 20,000 single-frame and 1,168 video based QA pairs to evaluate cognitive generalization and interpretability in both static and dynamic scenarios. Furthermore, we propose RailVQA-CoM, a collaborative large-small model framework that combines small-model efficiency with large-model cognition via a transparent three-module architecture and adaptive temporal sampling, improving perceptual generalization and enabling more efficient reasoning and planning. Experiments demonstrate that the proposed approach substantially improves performance, enhances interpretability, improves efficiency, and strengthens cross-domain generalization in autonomous driving systems. Code and datasets will be available at https://cybereye-bjtu.github.io/RailVQA.html.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。