arXiv:2603.24166cs.CV2026-03

用启发式推理先验提升少样本目标指代检测效率

Heuristic-inspired Reasoning Priors Facilitate Data-Efficient Referring Object Detection

  • 引入启发式先验指导候选框排序、预测融合与匹配
  • 在少样本条件下显著优于主流基线模型
  • 适合资源受限场景如机器人、AR等视觉语言任务

大多数指代对象检测(ROD)模型,尤其是现代定位检测器,设计用于数据丰富的环境,但在机器人、增强现实等实际部署中常面临严重标签稀缺问题。在此类场景下,端到端定位检测器需从零开始学习空间与语义结构,浪费宝贵样本。本文提出一个数据高效指代对象检测(De-ROD)任务,用于评估低数据和少样本设置下的性能。进一步提出HeROD框架,一种轻量级、模型无关的方法,将基于指代短语生成的可解释性启发式空间与语义先验注入现代DETR式流水线的三个阶段:候选框排序、预测融合与匈牙利匹配。通过引导训练与推理聚焦合理候选,提升标签效率与收敛性能。在RefCOCO、RefCOCO+和RefCOCOg数据集上,HeROD在标签稀缺条件下持续超越强基线模型。结果表明,融入简单可解释的推理先验,是实现更高效视觉语言理解的实用且可扩展路径。

原文摘要 · Abstract (English)

Most referring object detection (ROD) models, especially the modern grounding detectors, are designed for data-rich conditions, yet many practical deployments, such as robotics, augmented reality, and other specialized domains, would face severe label scarcity. In such regimes, end-to-end grounding detectors need to learn spatial and semantic structure from scratch, wasting precious samples. We ask a simple question: Can explicit reasoning priors help models learn more efficiently when data is scarce? To explore this, we first introduce a Data-efficient Referring Object Detection (De-ROD) task, which is a benchmark protocol for measuring ROD performance in low-data and few-shot settings. We then propose the HeROD (Heuristic-inspired ROD), a lightweight, model-agnostic framework that injects explicit, heuristic-inspired spatial and semantic reasoning priors, which are interpretable signals derived based on the referring phrase, into 3 stages of a modern DETR-style pipeline: proposal ranking, prediction fusion, and Hungarian matching. By biasing both training and inference toward plausible candidates, these priors promise to improve label efficiency and convergence performance. On RefCOCO, RefCOCO+, and RefCOCOg, HeROD consistently outperforms strong grounding baselines in scarce-label regimes. More broadly, our results suggest that integrating simple, interpretable reasoning priors provides a practical and extensible path toward better data-efficient vision-language understanding.

目标检测少样本学习视觉语言推理先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。