针对数据冗余问题,提出基于对称群的高效主动学习方法
Group-invariant Coresets for Data-efficient Active Learning

- 在变换群诱导的商空间中选择样本轨道,避免重复标注
- 在旋转图像数据上实现更高轨道覆盖率和更优标签效率
- 适用于具有明显对称性(如旋转、缩放)的数据集
主动学习通过查询最有信息量的未标注样本降低标注成本,但传统核集合方法忽略已知的数据对称性,可能在相同实例的不同变换版本上浪费预算。本文提出GRINCO,一种群不变核集合框架,其在由变换群诱导的商空间中进行采样,使选择操作基于轨道而非原始样本。该方法采用规范代表或学习到的轨道分离不变嵌入来定义实用的商空间度量,并将商空间k中心选择与通过轨道平均损失的不变训练相结合。我们进一步推导了一个泛化界,将超出轨道平均风险与商空间覆盖度、标签不确定性及轨道内变异性相关联。在具有尺度不变性的合成数据和具有旋转冗余的图像基准测试上,GRINCO提升了轨道覆盖度,并在群体诱导冗余显著时表现出比传统核集合基线更强的标签效率。
原文摘要 · Abstract (English)
Active learning reduces labeling cost by querying the most informative unlabeled samples, but standard coreset methods ignore known data symmetries and can waste budget on transformed versions of the same instance. We propose GRINCO, a group-invariant coreset framework that performs acquisition in the quotient space induced by a transformation group, so that selection operates on orbits rather than raw samples. The method uses either canonical representatives or learned orbit-separating invariant embeddings to define practical quotient metrics, and combines quotient-space k-center selection with invariant training through an orbit-averaged loss. We further derive a generalization bound that relates excess orbit-averaged risk to quotient-space coverage, label uncertainty, and intra-orbit variability. Experiments on synthetic scale-invariant data and image benchmarks with rotation-induced redundancy show that GRINCO improves orbit coverage and achieves stronger label efficiency than conventional coreset baselines, especially when group-induced redundancy is substantial.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。