用生成数据提升视觉定位效果,解决小样本难题。
Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding
- 通过盒外修复生成新图像,缓解标签错位问题。
- 综合难易度与过拟合风险,筛选最优训练数据。
- 在多个数据集和模型上表现稳定,适合数据少的场景。
视觉定位旨在根据文本查询定位图像区域。由于大规模数据构建困难,本文研究在数据稀缺条件下的有效学习方法。为应对数据不足,提出新颖框架POBF(Paint Outside the Box and Filter),通过盒外图像修复生成新数据,解决以往方法中的标签错位问题。同时,设计创新的数据筛选机制,结合难易度评分与过拟合评分,并由惩罚项平衡二者。在四个基准数据集上的实验表明,POBF相比仅使用真实数据的方法平均提升5.83%,优于领先基线2.29%–3.85%的准确率。此外,验证了其在不同生成模型、训练数据量及模型架构下的鲁棒性与泛化能力。
原文摘要 · Abstract (English)
Visual grounding aims to localize the image regions based on a textual query. Given the difficulty of large-scale data curation, we investigate how to effectively learn visual grounding under data-scarce settings in this paper. To address the data scarcity, we propose a novel framework, POBF (Paint Outside the Box and Filter). POBF synthesizes images by inpainting outside the box, tackling a label misalignment issue encountered in previous works. Furthermore, POBF leverages an innovative filtering scheme to select the most effective training data. This scheme combines a hardness score and an overfitting score, balanced by a penalty term. Extensive experiments across four benchmark datasets demonstrate that POBF consistently improves performance, achieving an average gain of 5.83\% over the real-data-only method and outperforming leading baselines by 2.29\%-3.85\% in accuracy. Additionally, we validate the robustness and generalizability of POBF across various generative models, training data sizes, and model architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。