针对稀有类别检索,提出高效主动学习策略,提升用户交互下的发现速度与准确率。
Positive-First Most Ambiguous: A Simple Active Learning Criterion for Interactive Retrieval of Rare Categories
- 优先选择可能为正例且处于分类边界附近的样本,兼顾效率与信息量。
- 在低预算、长尾数据下,显著提高早期检索准确率与类别覆盖度。
- 适合生态监测等需要快速识别罕见物种的交互式视觉检索场景。
现实世界中的细粒度视觉检索常需从大量未标注数据中以最少监督发现稀有概念,尤其在生物多样性监测、生态研究及长尾视觉领域尤为关键,目标类别可能仅占数据极小比例,导致严重类别不平衡。交互式检索结合相关性反馈提供可行方案:从少量查询开始,系统选择候选样本供用户二值标注,并迭代优化轻量级分类器。尽管主动学习(AL)常用于指导选择,但传统方法假设对称类先验和充足标注预算,在不平衡、低预算、低延迟环境下表现受限。本文提出正例优先最模糊(PF-MA)准则,显式应对类别不平衡的非对称性:优先选择靠近边界且可能为正例的样本,实现对细微视觉类别的快速发现并保持高信息量。相比标准方法过度采样负例,PF-MA始终返回高相关样本占比的小批量,显著提升早期检索效果与用户满意度。为衡量检索多样性,我们还引入类别覆盖度指标,评估所选正例对目标类视觉变异性覆盖程度。在长尾数据集(包括细粒度植物数据)上的实验表明,无论类别大小或描述符差异,PF-MA在覆盖率与分类性能上均持续优于强基线。结果表明,将主动学习与交互式细粒度检索中不对称且以用户为中心的目标对齐,可实现简单却强大的稀有与视觉微妙类别的检索方案。
原文摘要 · Abstract (English)
Real-world fine-grained visual retrieval often requires discovering a rare concept from large unlabeled collections with minimal supervision. This is especially critical in biodiversity monitoring, ecological studies, and long-tailed visual domains, where the target may represent only a tiny fraction of the data, creating highly imbalanced binary problems. Interactive retrieval with relevance feedback offers a practical solution: starting from a small query, the system selects candidates for binary user annotation and iteratively refines a lightweight classifier. While Active Learning (AL) is commonly used to guide selection, conventional AL assumes symmetric class priors and large annotation budgets, limiting effectiveness in imbalanced, low-budget, low-latency settings. We introduce Positive-First Most Ambiguous (PF-MA), a simple yet effective AL criterion that explicitly addresses the class imbalance asymmetry: it prioritizes near-boundary samples while favoring likely positives, enabling rapid discovery of subtle visual categories while maintaining informativeness. Unlike standard methods that oversample negatives, PF-MA consistently returns small batches with a high proportion of relevant samples, improving early retrieval and user satisfaction. To capture retrieval diversity, we also propose a class coverage metric that measures how well selected positives span the visual variability of the target class. Experiments on long-tailed datasets, including fine-grained botanical data, demonstrate that PF-MA consistently outperforms strong baselines in both coverage and classifier performance, across varying class sizes and descriptors. Our results highlight that aligning AL with the asymmetric and user-centric objectives of interactive fine-grained retrieval enables simple yet powerful solutions for retrieving rare and visually subtle categories in realistic human-in-the-loop settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。