arXiv:2603.10722cs.CVcs.AI2026-03

融合视觉与热成像,用交通规则知识提升无人机交通认知能力

UAV traffic scene understanding: A regulation embedded multi-modal network and a unified benchmark

  • 引入外部交通规则记忆库,用语义原型增强模型对复杂行为的理解
  • 在夜间和雾霾下仍保持高精度,热成像与可见光双向补偿提升鲁棒性
  • 首个光学-热红外统一基准数据集,含130万问答对,适合智能交通研究

从无人机平台进行交通场景理解对智能交通系统至关重要,因其部署灵活且覆盖范围广。然而,现有方法严重依赖可见光图像,在夜间、雾天等恶劣光照条件下性能显著下降。同时,现有视觉问答(VQA)模型仅能完成基础感知任务,缺乏评估复杂交通行为所需的领域特定法规知识。为此,我们提出一种新型多模态交通认知网络(MTCNet)。设计了原型引导的知识嵌入(PGKE)模块,利用外部交通规则记忆(TRM)中的高层语义原型,将领域知识锚定至视觉表征中,使模型能理解复杂交通行为并识别细微违规。此外,构建了质量感知光谱补偿(QASC)模块,利用可见光与热成像的互补特性实现双向上下文交互,有效补偿退化特征,确保复杂环境下的鲁棒表示。同时,构建了首个大规模光学-热红外基准数据集Traffic-VQA,包含8,180对齐图像及130万问答对,覆盖31种多样化场景。大量实验表明,MTCNet在认知与感知任务上均显著优于现有方法。数据集已开源:https://github.com/YuZhang-2004/UAV-traffic-scene-understanding。

原文摘要 · Abstract (English)

Traffic scene understanding from unmanned aerial vehicle (UAV) platforms is crucial for intelligent transportation systems due to its flexible deployment and wide-area monitoring capabilities. However, existing methods face significant challenges in real-world surveillance, as their heavy reliance on optical imagery leads to severe performance degradation under adverse illumination conditions like nighttime and fog. Furthermore, current Visual Question Answering (VQA) models are restricted to elementary perception tasks, lacking the domain-specific regulatory knowledge required to assess complex traffic behaviors. To address these limitations, we propose a novel Multi-modal Traffic Cognition Network (MTCNet) for robust UAV traffic scene understanding. Specifically, we design a Prototype-Guided Knowledge Embedding (PGKE) module that leverages high-level semantic prototypes from an external Traffic Regulation Memory (TRM) to anchor domain-specific knowledge into visual representations, enabling the model to comprehend complex behaviors and distinguish fine-grained traffic violations. Moreover, we develop a Quality-Aware Spectral Compensation (QASC) module that exploits the complementary characteristics of optical and thermal modalities to perform bidirectional context exchange, effectively compensating for degraded features to ensure robust representation in complex environments. In addition, we construct Traffic-VQA, the first large-scale optical-thermal infrared benchmark for cognitive UAV traffic understanding, comprising 8,180 aligned image pairs and 1.3 million question-answer pairs across 31 diverse types. Extensive experiments demonstrate that MTCNet significantly outperforms state-of-the-art methods in both cognition and perception scenarios. The dataset is available at https://github.com/YuZhang-2004/UAV-traffic-scene-understanding.

无人机交通多模态视觉问答规则嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。