通过迭代重对齐提升图像分类的鲁棒性,无需依赖理想数据
Robust Canonicalization through Bootstrapped Data Re-Alignment
- 用迭代方法逐步减少样本差异,恢复数据对齐假设
- 在4个细粒度识别任务中超越等变与传统归一化模型
- 适合数据有方向、尺度偏差的图像分类场景
细粒度视觉分类(FGVC)任务如昆虫和鸟类识别,需敏感捕捉细微视觉特征,同时抵抗空间变换影响。现有方法依赖大量数据增强(需强模型)或等变结构(限制表达能力且增加开销)。归一化方法可屏蔽此类偏差,但通常依赖训练数据已对齐的先验假设,而真实数据往往不满足此条件,导致归一化器脆弱。本文提出一种自举算法,通过迭代重对齐训练样本,逐步降低方差并恢复对齐假设。我们在任意紧群下建立收敛性保证,并在4个FGVC基准上验证:该方法始终优于等变与归一化基线,性能与数据增强相当。
原文摘要 · Abstract (English)
Fine-grained visual classification (FGVC) tasks, such as insect and bird identification, demand sensitivity to subtle visual cues while remaining robust to spatial transformations. A key challenge is handling geometric biases and noise, such as different orientations and scales of objects. Existing remedies rely on heavy data augmentation, which demands powerful models, or on equivariant architectures, which constrain expressivity and add cost. Canonicalization offers an alternative by shielding such biases from the downstream model. In practice, such functions are often obtained using canonicalization priors, which assume aligned training data. Unfortunately, real-world datasets never fulfill this assumption, causing the obtained canonicalizer to be brittle. We propose a bootstrapping algorithm that iteratively re-aligns training samples by progressively reducing variance and recovering the alignment assumption. We establish convergence guarantees under mild conditions for arbitrary compact groups, and show on four FGVC benchmarks that our method consistently outperforms equivariant, and canonicalization baselines while performing on par with augmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。