通过逻辑正则化提升视觉分类模型泛化能力
Interpret Your Decision: Logical Reasoning Regularization for Generalization in Visual Classification
- 引入L-Reg逻辑正则化,约束特征分布与分类器权重
- 在多域泛化和新类别发现任务中显著提升性能
- 增强模型可解释性,聚焦关键特征如人脸识别人物
视觉模型在图像分类上表现优异,但在面对未见过的数据时泛化能力不足,例如跨域分类或发现新类别。本文探讨逻辑推理与深度学习泛化之间的关系,提出一种名为L-Reg的逻辑正则化方法,将逻辑分析框架引入图像分类。L-Reg通过降低模型在特征分布和分类器权重上的复杂度,提升泛化性能。理论分析与实验表明,该方法在多域泛化与广义类别发现任务中均有效。在包含未知类别和未见域的复杂真实场景中,L-Reg持续改善泛化表现,验证其实际有效性。
原文摘要 · Abstract (English)
Vision models excel in image classification but struggle to generalize to unseen data, such as classifying images from unseen domains or discovering novel categories. In this paper, we explore the relationship between logical reasoning and deep learning generalization in visual classification. A logical regularization termed L-Reg is derived which bridges a logical analysis framework to image classification. Our work reveals that L-Reg reduces the complexity of the model in terms of the feature distribution and classifier weights. Specifically, we unveil the interpretability brought by L-Reg, as it enables the model to extract the salient features, such as faces to persons, for classification. Theoretical analysis and experiments demonstrate that L-Reg enhances generalization across various scenarios, including multi-domain generalization and generalized category discovery. In complex real-world scenarios where images span unknown classes and unseen domains, L-Reg consistently improves generalization, highlighting its practical efficacy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。