构建铁路入侵感知新基准,提升视觉模型对潜在风险的时空推理能力
CogRail: Benchmarking VLMs in Cognitive Intrusion Perception for Intelligent Railway Transportation Systems
- 设计融合时空上下文的问答标注数据集,支持认知级入侵感知评测
- 现有大模型在复杂时空推理上表现不佳,准确率不足60%
- 提出多任务联合微调框架,显著提升安全关键场景下的推理准确性
精准且早期地感知潜在入侵目标对保障铁路运输安全至关重要。然而,现有系统多局限于固定视野内的物体分类,并依赖规则启发式判断入侵状态,常忽略具有潜在入侵风险的目标。预判此类风险需对目标对象(OOI)的空间上下文和时间动态具备认知理解,这对传统视觉模型构成挑战。为此,我们提出新基准CogRail,整合开源数据集与认知驱动的问答标注,支持时空推理与预测。基于此基准,我们系统评估了主流视觉语言模型(VLMs)在多模态提示下的表现,揭示其在该领域的优劣。进一步通过微调提升性能,提出融合位置感知、运动预测与威胁分析三任务的联合微调框架,实现通用基础模型向专用认知入侵感知模型的有效迁移。大量实验表明,当前大规模多模态模型在认知入侵感知任务中难以胜任复杂的时空推理,准确率普遍低于60%。而所提联合微调框架显著提升模型表现,凸显结构化多任务学习在提高准确率与可解释性方面的优势。代码将公开于 https://github.com/Hub-Tian/CogRail。
原文摘要 · Abstract (English)
Accurate and early perception of potential intrusion targets is essential for ensuring the safety of railway transportation systems. However, most existing systems focus narrowly on object classification within fixed visual scopes and apply rule-based heuristics to determine intrusion status, often overlooking targets that pose latent intrusion risks. Anticipating such risks requires the cognition of spatial context and temporal dynamics for the object of interest (OOI), which presents challenges for conventional visual models. To facilitate deep intrusion perception, we introduce a novel benchmark, CogRail, which integrates curated open-source datasets with cognitively driven question-answer annotations to support spatio-temporal reasoning and prediction. Building upon this benchmark, we conduct a systematic evaluation of state-of-the-art visual-language models (VLMs) using multimodal prompts to identify their strengths and limitations in this domain. Furthermore, we fine-tune VLMs for better performance and propose a joint fine-tuning framework that integrates three core tasks, position perception, movement prediction, and threat analysis, facilitating effective adaptation of general-purpose foundation models into specialized models tailored for cognitive intrusion perception. Extensive experiments reveal that current large-scale multimodal models struggle with the complex spatial-temporal reasoning required by the cognitive intrusion perception task, underscoring the limitations of existing foundation models in this safety-critical domain. In contrast, our proposed joint fine-tuning framework significantly enhances model performance by enabling targeted adaptation to domain-specific reasoning demands, highlighting the advantages of structured multi-task learning in improving both accuracy and interpretability. Code will be available at https://github.com/Hub-Tian/CogRail.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。