arXiv:2609.09025cs.CV2026-09

让视觉模型像人一样聚焦关键区域,提升目标检测精度。

Task-driven Processing with Coarse-to-Fine Glimpse-based Active Perception

论文配图:Task-driven Processing with Coarse-to-Fine Glimpse-based Active Perception
图 1 · 摘自论文原文
  • 分步聚焦场景关键区域,用任务信息引导高分辨率处理。
  • 在多个数据集上将检测准确率提升最高达20%。
  • 适合资源受限场景下追求高精度的检测任务。

当前主流视觉模型对图像进行全图处理,缺乏选择性关注关键区域的能力。在实例检测等任务中,尤其在高分辨率、杂乱场景下,重要细节常因图像缩放和计算限制而丢失。本文提出粗到精注视式主动感知(CF-GAP),作为任务驱动的前端模块,增强现有实例检测器的高分辨率处理能力。CF-GAP根据任务信息,依次投射有限视野的注视片段,迭代聚焦最相关区域,并以高分辨率交由下游检测器处理。通过避免全图计算并消除无关干扰,CF-GAP在HR-InsDet与Robotools基准上,使多种先进检测器的平均精度(AP)提升最高达20%,同时让轻量级检测器超越其大型版本。

原文摘要 · Abstract (English)

State-of-the-art vision models process images in their entirety, lacking the ability to selectively zoom in on relevant regions. This limitation is particularly acute in scenarios where processing must be conditioned on a specific task - such as instance detection, which requires localizing a specific object in a high-resolution, cluttered scene. In such settings, critical details are easily lost as images are often resized to match the model dimensions and computational constraints. We introduce Coarse-to-Fine Glimpse-based Active Perception (CF-GAP), a task-driven front-end that enhances high-resolution processing of existing instance detectors. CF-GAP selectively directs a sequence of limited view glimpses across the scene, utilizing task information to iteratively refine focus on the most relevant regions. These localized regions are then processed at high resolution by a downstream instance detector. By avoiding full-image processing and eliminating irrelevant confounding information, CF-GAP improves Average Precision (AP) by up to 20% across various state-of-the-art instance detectors on the HR-InsDet and Robotools benchmarks, while further enabling lightweight detectors to outperform their larger counterparts.

目标检测主动感知高分辨率效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。