arXiv:2505.24372cs.CV2025-05被引 1

通过分布感知标注提升视觉定位数据质量,避免噪声与冗余。

Beyond Quantity: Distribution-Aware Labeling for Visual Grounding

  • 双路径标注:封闭集保可靠,开放集扩词汇,引入新概念。
  • 显式拓展分布外表达,覆盖更广语义范围。
  • 过滤噪声与冗余样本,均衡语言与视觉内容分布。

视觉定位需要大量且多样的区域-文本配对数据。然而,人工标注成本高,固定词表限制可扩展性与泛化能力。现有伪标签流程常对偏差分布过拟合,生成噪声或冗余样本。通过对数据质量与分布覆盖的系统分析,我们发现性能提升更多源于有效分布扩展,而非单纯的数据量增加。为此,提出DAL(分布感知标注)框架:首先采用双驱动标注模块,封闭集路径提供可靠伪标签,开放集路径丰富词汇并引入新概念;同时显式进行分布外(OOD)表达扩展,拓宽语义覆盖。随后设计一致性与分布感知过滤模块,剔除噪声和冗余区域-文本对,重新平衡语言与视觉内容的分布,从而提升数据质量与训练效率。在三个基准上的大量实验表明,该方法持续优于强基线,达到当前最优效果,凸显分布感知标注在构建可扩展、鲁棒视觉定位数据集中的关键作用。

原文摘要 · Abstract (English)

Visual grounding requires large and diverse region-text pairs. However, manual annotation is costly and fixed vocabularies restrict scalability and generalization. Existing pseudo-labeling pipelines often overfit to biased distributions and generate noisy or redundant samples. Through our systematic analysis of data quality and distributional coverage, we find that performance gains come less from raw data volume and more from effective distribution expansion. Motivated by this insight, we propose DAL, a distribution-aware labeling framework for visual grounding. The proposed method first employs a dual-driven annotation module, where a closed-set path provides reliable pseudo labels and an open-set path enriches vocabulary and introduces novel concepts; meanwhile, it further performs explicit out-of-distribution (OOD) expression expansion to broaden semantic coverage. We then propose a consistency- and distribution-aware filtering module to discard noisy or redundant region-text pairs and rebalance underrepresented linguistic and visual content, thereby improving both data quality and training efficiency. Extensive experiments on three benchmarks demonstrate that our method consistently outperforms strong baselines and achieves state-of-the-art results, underscoring the critical role of distribution-aware labeling in building scalable and robust visual grounding datasets.

视觉定位伪标签数据质量分布建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。