arXiv:2608.30621cs.CVcs.AI2026-08中稿 · EMNLP

用自动生成的图文对,高效挑选难标注的模糊图像进行主动学习。

Cost-efficient Active Learning for Referring Image Segmentation and Grounding

论文配图:Cost-efficient Active Learning for Referring Image Segmentation and Grounding
图 1 · 摘自论文原文
  • 基于基础模型生成辅助图文对,识别视觉模糊区域
  • 提出置信度分散度指标,优先选择竞争性强的图像
  • 界面优化使标注速度提升1.6倍,适合资源受限场景

在视觉定位(VG)中,获取自然语言指代表达与区域标注(如掩码或边界框)是主要瓶颈,因为标注者需编写能区分目标区域与视觉相似区域的描述。本文在仅提供原始图像、无配套文本的现实场景下,提出面向VG的主动学习(AL)方法。由于缺乏真实文本,样本选择需估计哪些图像包含需要区分性指代表达的模糊区域。为此,我们利用基础模型生成辅助区域-文本对,并引入「被指区域模糊性」(Referred Region Ambiguity)这一新采集函数,衡量模型置信度是否集中于单一区域或分散于多个候选区域。该指标可优先选择跨区域竞争强烈的图像,因其视觉模糊性带来更高信息量。我们还设计了指代表达标注界面,通过少量点击帮助标注者快速聚焦生成区分性语言。在RIS和REC基准上的实验表明,本方法持续优于多个主流主动学习基线;用户研究显示,标注效率最高提升1.6倍。

原文摘要 · Abstract (English)

Collecting natural-language referring expressions along with region annotations, such as masks or boxes, is a major bottleneck in visual grounding (VG), as annotators must write descriptions that distinguish target regions from visually similar ones. We tackle this by formulating active learning (AL) for VG under the realistic setting where only raw images are available without accompanying text. Since ground-truth text is unavailable, sample selection must estimate which images contain ambiguous regions that would require discriminative referring expressions. To address this, we generate auxiliary region-text pairs using foundation models, and introduce Referred Region Ambiguity, a new acquisition function that measures whether the model's confidence collapses onto a single region or disperses across multiple candidates. It allows our method to prioritize images with strong cross-region competition, which are more informative due to their visual ambiguity. We also design a referring-expression annotation interface that helps annotators quickly focus on writing discriminative language with a few clicks. Experiments on RIS and REC benchmarks show that our AL framework consistently outperforms several AL baselines, while a user study shows up to 1.6X faster description labeling of ours.

主动学习视觉定位标注效率基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。