用风格迁移构建24个图像数据集,研究模型在分布偏移下的鲁棒性。
Stylized Meta-Album: Group-bias injection with style transfer to study robustness against distribution shifts
- 通过风格迁移生成12组内容与12组风格化数据,共4800个分组。
- 提升群体多样性后,公平性指标显著变化,算法排名重新排序。
- 适用于评估模型在复杂分布偏移下的鲁棒性与公平性,适合研究者使用。
我们提出Stylized Meta-Album(SMA),一个包含24个数据集(12个内容数据集和12个风格化数据集)的图像分类元数据集,旨在推动分布外(OOD)泛化及相关研究。SMA基于12个主体分类数据集,通过风格迁移技术生成,涵盖多种主题(物体、植物、动物、人类动作、纹理)与多重风格,共4800个分组。该设计支持对组别与类别的灵活控制,可配置为反映多样化的基准场景。虽然真实数据收集难以覆盖广泛分组,但SMA通过灵活调整风格、主题类别与领域,实现大规模且可配置的分组结构,扩展了组与类的多样性,并开辟新方法论方向,用于评估模型在多数少数群体、不同组不平衡及复杂领域偏移下的表现,以及公平性、鲁棒性与适应性。我们构建了两个基准:(1) 新型OOD泛化与组公平性基准,利用其多样性评估现有方法;结果表明,尽管简单平衡与利用组信息的算法仍具竞争力,但增加组多样性会显著影响公平性,改变算法相对性能。(2) 无监督域适应(UDA)基准,利用组多样性评估更多场景,相比现有工作,在闭集设置与UniDA设置下误差范围分别降低73%与28%。这些应用凸显SMA对传统基准结果的显著影响。
原文摘要 · Abstract (English)
We introduce Stylized Meta-Album (SMA), a new image classification meta-dataset comprising 24 datasets (12 content datasets, and 12 stylized datasets), designed to advance studies on out-of-distribution (OOD) generalization and related topics. Created using style transfer techniques from 12 subject classification datasets, SMA provides a diverse and extensive set of 4800 groups, combining various subjects (objects, plants, animals, human actions, textures) with multiple styles. SMA enables flexible control over groups and classes, allowing us to configure datasets to reflect diverse benchmark scenarios. While ideally, data collection would capture extensive group diversity, practical constraints often make this infeasible. SMA addresses this by enabling large and configurable group structures through flexible control over styles, subject classes, and domains-allowing datasets to reflect a wide range of real-world benchmark scenarios. This design not only expands group and class diversity, but also opens new methodological directions for evaluating model performance across diverse group and domain configurations-including scenarios with many minority groups, varying group imbalance, and complex domain shifts-and for studying fairness, robustness, and adaptation under a broader range of realistic conditions. To demonstrate SMA's effectiveness, we implemented two benchmarks: (1) a novel OOD generalization and group fairness benchmark leveraging SMA's domain, class, and group diversity to evaluate existing benchmarks. Our findings reveal that while simple balancing and algorithms utilizing group information remain competitive as claimed in previous benchmarks, increasing group diversity significantly impacts fairness, altering the superiority and relative rankings of algorithms. We also propose to use \textit{Top-M worst group accuracy} as a new hyperparameter tuning metric, demonstrating broader fairness during optimization and delivering better final worst-group accuracy for larger group diversity. (2) An unsupervised domain adaptation (UDA) benchmark utilizing SMA's group diversity to evaluate UDA algorithms across more scenarios, offering a more comprehensive benchmark with lower error bars (reduced by 73\% and 28\% in closed-set setting and UniDA setting, respectively) compared to existing efforts. These use cases highlight SMA's potential to significantly impact the outcomes of conventional benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。