arXiv:2510.23785cs.CVcs.AI2025-10中稿 · the 2026 IEEE 2nd …

用视觉基础模型提升无类别物体计数的结构一致性

CountFormer: A Transformer Framework for Learning Visual Repetition and Structure in Class-Agnostic Object Counting

  • 基于DINOv2提取图像特征,结合位置编码与轻量解码器生成密度图
  • 在FSC-147上MAE达19.06,对复杂结构物体过计数减少
  • 适合关注表征质量对无样本计数影响的研究者

人类可通过观察视觉重复与构型计数陌生物体,而非依赖类别。但许多无样本计数模型在此类场景下表现不佳,尤其在对称部件、重复子结构或部分遮挡时易过计。本文提出CountFormer,将密度回归框架CounTR中的图像编码器替换为自监督视觉基础模型DINOv2,融合二维显式位置嵌入,由轻量卷积网络解码生成密度图,其积分得最终计数。目标并非设计新架构,而是探究基础模型表征在严格无样本设定下的结构一致性提升效果。在FSC-147基准上,CountFormer取得竞争性结果(MAE 19.06,RMSE 118.45)。定性分析显示,对部分结构复杂物体的部件级过计错误减少,整体误差与先前方法基本持平。敏感性分析表明,评估指标受少数极高密度场景显著影响。结果凸显表征质量在无样本物体计数中的关键作用。

原文摘要 · Abstract (English)

Humans can often count unfamiliar objects by observing visual repetition and composition, rather than relying only on object categories. However, many exemplar-free counting models struggle in such situations and may overcount when objects contain symmetric components, repeated substructures, or partial occlusion. We introduce CountFormer, a controlled adaptation of a density-regression framework inspired by CounTR, where the image encoder is replaced with the self-supervised vision foundation model DINOv2. The resulting transformer features are combined with explicit two-dimensional positional embeddings and decoded by a lightweight convolutional network to produce a density map whose integral gives the final count. Our goal is not to propose a new counting architecture, but to study whether foundation-based representations improve structural consistency under a strictly exemplar-free setting. On FSC-147, CountFormer achieves competitive performance under the official benchmark (MAE 19.06, RMSE 118.45). Qualitative analysis suggests fewer part-level overcounting errors for some structurally complex objects, while overall error remains broadly consistent with prior approaches. Sensitivity analysis shows that evaluation metrics are strongly affected by a small number of extreme high-density scenes. Overall, the results highlight the role of representation quality in exemplar-free object counting.

物体计数视觉基础模型结构一致性Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。