数据不平衡未必坏事,模型对群体分布敏感度取决于隐空间分离程度。
Representation Invariance and Allocation: When Subgroup Balance Matters
- 通过隐空间分离度分析模型对子群体数据的依赖性
- 发现部分子群体缺失时性能不降反升,与平衡数据假设相反
- 为大模型微调阶段的数据收集提供可量化的决策依据
训练数据中不同人口群体的不均衡分布影响模型在不同人群间的泛化能力。传统做法认为平衡子群体数据能优化整体性能,但近期实证结果表明:在某些情况下,数据分布不均反而提升子群体表现;而在另一些情形下,训练中完全缺失某个子群体也未影响其性能。本文系统研究了四种视觉-语言模型在不同数据组成下的子群体表现敏感性,提出隐空间分离假设:部分微调模型对子群体数据的依赖程度由预训练模型中子群体在隐空间的分离程度决定。我们对该假设进行形式化、理论推导并实证验证。最后,展示该方法在基础模型微调中的实际应用,证明通过量化隐空间子群体分离度,可指导数据收集与平衡策略制定。
原文摘要 · Abstract (English)
Unequal representation of demographic groups in training data poses challenges to model generalisation across populations. Standard practice assumes that balancing subgroup representation optimises performance. However, recent empirical results contradict this assumption: in some cases, imbalanced data distributions actually improve subgroup performance, while in others, subgroup performance remains unaffected by the absence of an entire subgroup during training. We conduct a systematic study of subgroup allocation across four vision and language models, varying training data composition to characterise the sensitivity of subgroup performance to data balance. We propose the latent separation hypothesis, which states that a partially fine-tuned model's dependence on subgroup representation is determined by the degree of separation between subgroups in the latent space of the pre-trained model. We formalise this hypothesis, provide theoretical analysis, and validate it empirically. Finally, we present a practical application to foundation model fine-tuning, demonstrating that quantitative analysis of latent subgroup separation can inform data collection and balancing decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。