arXiv:2511.02667cs.LGcs.AI2025-11NeurIPS被引 3

提出高效评估框架与新模型,显著提升组合泛化能力。

Scalable Evaluation and Neural Models for Compositional Generalization

  • 构建统一评估框架,计算量从组合级降为常数级
  • 在5000+模型上验证,新模型准确率提升23.43%
  • 新模型参数开销仅16%,适合大规模应用

组合泛化是现代机器学习的关键挑战,要求模型能预测已知概念的未知组合。然而,当前评估缺乏标准化协议,现有基准往往重效率轻严谨性。同时,通用视觉架构缺少必要归纳偏置,现有增强方法又牺牲可扩展性。为此,本文提出:1)统一并扩展先前方法的严格评估框架,将计算需求从组合级降至常数级;2)对监督视觉主干网络的组合泛化现状进行大规模评估,训练超过5000个模型;3)提出属性不变网络(Attribute Invariant Networks),在组合泛化上达到新帕累托前沿,相比基线提升23.43%准确率,参数开销从600%降至16%。代码已开源。

原文摘要 · Abstract (English)

Compositional generalization-a key open challenge in modern machine learning-requires models to predict unknown combinations of known concepts. However, assessing compositional generalization remains a fundamental challenge due to the lack of standardized evaluation protocols and the limitations of current benchmarks, which often favor efficiency over rigor. At the same time, general-purpose vision architectures lack the necessary inductive biases, and existing approaches to endow them compromise scalability. As a remedy, this paper introduces: 1) a rigorous evaluation framework that unifies and extends previous approaches while reducing computational requirements from combinatorial to constant; 2) an extensive and modern evaluation on the status of compositional generalization in supervised vision backbones, training more than 5000 models; 3) Attribute Invariant Networks, a class of models establishing a new Pareto frontier in compositional generalization, achieving a 23.43% accuracy improvement over baselines while reducing parameter overhead from 600% to 16% compared to fully disentangled counterparts. Our code is available at https://github.com/IBM/scalable-compositional-generalization.

组合泛化评估框架视觉模型可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。