arXiv:2603.16129cs.CV2026-03中稿 · CVPR被引 4

让AI看图数数更准,还能理解物体位置和数量关系。

Boosting Quantitive and Spatial Awareness for Zero-Shot Object Counting

  • 用数值提示引导视觉语言模型,提升数量感知能力。
  • 在FSC-147上实现新高精度,跨数据集零样本测试表现优异。
  • 适合需要不依赖训练样本的智能计数场景,如监控、仓储。

零样本目标计数(ZSOC)旨在无需视觉样例的情况下,根据文本描述统计任意类别物体的数量。然而,现有方法常将计数视为粗粒度检索任务,缺乏精细的数量感知能力,且因模型适配时特征空间畸变,导致空间敏感性差、泛化性能下降。为此,本文提出QICA框架,融合数量感知与鲁棒的空间聚合机制。设计协同提示策略(SPS),通过数值条件提示对视觉与语言编码器进行适配,弥合语义识别与量化推理之间的鸿沟。提出代价聚合解码器(CAD),直接在视觉-文本相似性图上操作,通过空间聚合优化该图,防止过拟合并保持零样本可迁移性。引入多层级数量对齐损失($/mathcal{L}_{MQA}$),确保整个流程中数量一致性。在FSC-147上实验结果具有竞争力,于CARPK与ShanghaiTech-A上的零样本评估验证了对未见领域的优越泛化能力。

原文摘要 · Abstract (English)

Zero-shot object counting (ZSOC) aims to enumerate objects of arbitrary categories specified by text descriptions without requiring visual exemplars. However, existing methods often treat counting as a coarse retrieval task, suffering from a lack of fine-grained quantity awareness. Furthermore, they frequently exhibit spatial insensitivity and degraded generalization due to feature space distortion during model adaptation.To address these challenges, we present \textbf{QICA}, a novel framework that synergizes \underline{q}uantity percept\underline{i}on with robust spatial \underline{c}ast \underline{a}ggregation. Specifically, we introduce a Synergistic Prompting Strategy (\textbf{SPS}) that adapts vision and language encoders through numerically conditioned prompts, bridging the gap between semantic recognition and quantitative reasoning. To mitigate feature distortion, we propose a Cost Aggregation Decoder (\textbf{CAD}) that operates directly on vision-text similarity maps. By refining these maps through spatial aggregation, CAD prevents overfitting while preserving zero-shot transferability. Additionally, a multi-level quantity alignment loss ($\mathcal{L}_{MQA}$) is employed to enforce numerical consistency across the entire pipeline. Extensive experiments on FSC-147 demonstrate competitive performance, while zero-shot evaluation on CARPK and ShanghaiTech-A validates superior generalization to unseen domains.

目标计数零样本视觉语言空间感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。