arXiv:2507.07102cs.LG2025-07ICML被引 16

数据多样性比数量更重要,决定视觉组合泛化能力。

Does Data Scaling Lead to Visual Compositional Generalization?

  • 用可控实验验证数据多样性对组合泛化的关键作用
  • 组合覆盖度提升使模型学会线性分解概念结构
  • 适合关注高效组合学习与数据构建的研究者

组合理解是人类智能的核心,但当前视觉模型是否具备尚不明确。主流机器学习范式认为扩大数据和模型规模可提升分布外性能,包括组合泛化。我们通过受控实验系统地改变数据规模、概念多样性和组合覆盖率进行验证。结果表明,组合泛化由数据多样性驱动,而非单纯的数据规模。更高的组合覆盖率迫使模型发现线性因子化的表征结构,使概念可分解为加性成分。我们证明该结构是效率的关键,可从少量观测组合实现完美泛化。评估预训练模型(DINO、CLIP)发现其表现高于随机水平但不完美,表明该结构部分存在。研究呼吁更重视构建多样化数据集,并关注促进高效组合学习的表征结构。代码见:https://github.com/oshapio/visual-compositional-generalization。

原文摘要 · Abstract (English)

Compositional understanding is crucial for human intelligence, yet it remains unclear whether contemporary vision models exhibit it. The dominant machine learning paradigm is built on the premise that scaling data and model sizes will improve out-of-distribution performance, including compositional generalization. We test this premise through controlled experiments that systematically vary data scale, concept diversity, and combination coverage. We find that compositional generalization is driven by data diversity, not mere data scale. Increased combinatorial coverage forces models to discover a linearly factored representational structure, where concepts decompose into additive components. We prove this structure is key to efficiency, enabling perfect generalization from few observed combinations. Evaluating pretrained models (DINO, CLIP), we find above-random yet imperfect performance, suggesting partial presence of this structure. Our work motivates stronger emphasis on constructing diverse datasets for compositional generalization, and considering the importance of representational structure that enables efficient compositional learning. Code available at https://github.com/oshapio/visual-compositional-generalization.

组合泛化数据多样性表征结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。