arXiv:2410.08258cs.CV2024-10被引 26

构建风格严格不同的新数据集,揭示大模型泛化能力的虚假假象。

In Search of Forgotten Domain Generalization

  • 从LAION中提取自然与渲染风格数据,确保与ImageNet等测试集风格迥异
  • 发现模型性能主要依赖域内样本,而非真正跨域泛化
  • 提出最优混合比例,助力实现可量化的跨域鲁棒性评估

域外(OOD)泛化指模型在已知域上训练后对未见域的适应能力。图像识别时代,评测集设计为严格不同于训练域的风格。但随着基础模型和大规模网络数据集的兴起,数据覆盖范围广泛,导致测试集存在域污染风险。本文创建了从LAION中子采样的两大数据集:LAION-Natural与LAION-Rendition,其风格严格区别于ImageNet和DomainNet的测试集。在这些数据上训练CLIP模型发现,其性能的显著部分由域内样本解释,表明图像时代面临的OOD挑战依然存在,而基于网络数据的训练仅制造了泛化假象。通过系统探索自然与渲染数据的混合比例,我们识别出提升跨域泛化效果的最佳组合。本研究提供的数据集与结果重新建立了大规模下有意义的OOD鲁棒性评估体系,是提升模型鲁棒性的关键前提。

原文摘要 · Abstract (English)

Out-of-Domain (OOD) generalization is the ability of a model trained on one or more domains to generalize to unseen domains. In the ImageNet era of computer vision, evaluation sets for measuring a model's OOD performance were designed to be strictly OOD with respect to style. However, the emergence of foundation models and expansive web-scale datasets has obfuscated this evaluation process, as datasets cover a broad range of domains and risk test domain contamination. In search of the forgotten domain generalization, we create large-scale datasets subsampled from LAION -- LAION-Natural and LAION-Rendition -- that are strictly OOD to corresponding ImageNet and DomainNet test sets in terms of style. Training CLIP models on these datasets reveals that a significant portion of their performance is explained by in-domain examples. This indicates that the OOD generalization challenges from the ImageNet era still prevail and that training on web-scale data merely creates the illusion of OOD generalization. Furthermore, through a systematic exploration of combining natural and rendition datasets in varying proportions, we identify optimal mixing ratios for model generalization across these domains. Our datasets and results re-enable meaningful assessment of OOD robustness at scale -- a crucial prerequisite for improving model robustness.

域泛化鲁棒性评估CLIP数据污染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。