arXiv:2601.11322cs.CVcs.LO2026-01中稿 · publication in IEE…被引 1

用逻辑推理提升视觉语言模型对关键事件的识别能力

Enhancing Vision Language Models with Logic Reasoning for Situational Awareness

  • 通过显式逻辑推理融合传统视觉方法与视觉语言模型
  • 智能微调策略显著提高识别准确率,优于随机选择
  • 推理阶段生成输出理由,帮助判断结果可靠性

视觉语言模型(VLMs)能从图像和视频中生成高层、可解释的复杂活动描述,适用于情境感知(SA)场景。在这些场景中,重点是高可靠性和准确性地识别罕见但重要的事件,同时提取细粒度细节并评估识别质量。本文提出一种将VLMs与传统计算机视觉方法通过显式逻辑推理结合的方法,从三个方面增强情境感知:(a) 提取细粒度事件细节,(b) 采用智能微调(FT)策略,显著优于无信息选择,(c) 在推理过程中为VLM输出生成解释。实验证明,该智能微调机制提升了准确率,并在推理阶段提供确认输出有效性或揭示其可疑性的有效手段。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) offer the ability to generate high-level, interpretable descriptions of complex activities from images and videos, making them valuable for situational awareness (SA) applications. In such settings, the focus is on identifying infrequent but significant events with high reliability and accuracy, while also extracting fine-grained details and assessing recognition quality. In this paper, we propose an approach that integrates VLMs with traditional computer vision methods through explicit logic reasoning to enhance SA in three key ways: (a) extracting fine-grained event details, (b) employing an intelligent fine-tuning (FT) strategy that achieves substantially higher accuracy than uninformed selection, and (c) generating justifications for VLM outputs during inference. We demonstrate that our intelligent FT mechanism improves the accuracy and provides a valuable means, during inferencing, to either confirm the validity of the VLM output or indicate why it may be questionable.

视觉语言模型逻辑推理情境感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。