arXiv:2607.05859cs.CV2026-07

让视觉语言模型像人一样有选择地查看工地细节,提升效率与可靠性。

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring

论文配图:AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring
图 1 · 摘自论文原文
  • 先用低分辨率全局图粗略判断,仅在必要时调用高分辨率局部图像。
  • 在远距离和低分辨率条件下,准确率提升且视觉令牌使用减少60%以上。
  • 适合需要高效、鲁棒的野外工地监控场景,尤其关注资源优化的部署者。

视觉语言模型(VLMs)在工地监控中具有潜力,现有针对建筑领域的VLM主要通过单一全局图像的问答式微调进行改进。然而,这种直接方法在真实环境部署中仍受限于操作范围、低分辨率输入下的可靠性以及推理效率。为此,我们提出AVA-VLM,一种模仿人类从粗到细推理策略的自适应视觉注意力-视觉语言模型。AVA-VLM首先对低分辨率全局图像进行推理,仅在需要详细检查时才请求高分辨率局部图像,类似人类巡视时聚焦关键区域。我们还构建了一个区域感知的思维链数据集,指导模型何时检测、何处裁剪、如何利用局部证据。实验表明,AVA-VLM在长距离与低分辨率条件下显著提升可靠性,并大幅降低视觉令牌使用量。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) are promising for construction-site monitoring, and recent construction-tailored VLMs have primarily adapted pretrained VLMs through direct QA-style fine-tuning from a single global image. We argue that this direct paradigm remains limited for in-the-wild deployment in terms of operational range, reliability under reduced-resolution inputs, and inference efficiency. To address these challenges, we propose AVA-VLM, an Adaptive Visual Attention-Vision Language Model that follows a human-inspired coarse-to-fine reasoning strategy. AVA-VLM first reasons over a low-resolution global image and selectively requests a high-resolution local crop only when detailed inspection is needed, similar to how a human inspector zooms in on hard-to-see yet important areas. We further introduce a region-aware Chain-of-Thought dataset that teaches the model when to inspect, where to crop, and how to use local evidence. Experiments show that AVA-VLM improves reliability under long-distance and reduced-resolution conditions while substantially reducing visual-token usage.

视觉语言模型工地监控注意力机制推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。