arXiv:2410.06303cs.LGcs.AI2024-10被引 14

提出新方法应对训练中缺失的属性组合,让模型更像人一样推理。

Compositional Risk Minimization

  • 用能量模型分解属性,构建可组合的损失函数
  • 在多个基准数据集上显著提升对未见属性组合的泛化能力
  • 适合需要强泛化能力的复杂场景建模任务

组合泛化是实现数据高效智能机器的关键一步,使其能像人类一样进行泛化。本文针对一种挑战性分布偏移——组合偏移,即某些属性组合在训练中完全缺失但在测试分布中出现。这种偏移考验模型对新属性组合的判别式泛化能力。我们采用灵活的加性能量分布建模数据,每个能量项代表一个属性,并推导出一种替代经验风险最小化的简单方法,称为组合风险最小化(CRM)。先训练一个加性能量分类器预测多个属性,再调整该分类器以应对组合偏移。我们对CRM进行了广泛的理论分析,证明其可外推至已见属性组合的仿射包络。在基准数据集上的实证评估表明,与文献中针对各类子群体偏移设计的方法相比,CRM展现出更强的鲁棒性。

原文摘要 · Abstract (English)

Compositional generalization is a crucial step towards developing data-efficient intelligent machines that generalize in human-like ways. In this work, we tackle a challenging form of distribution shift, termed compositional shift, where some attribute combinations are completely absent at training but present in the test distribution. This shift tests the model's ability to generalize compositionally to novel attribute combinations in discriminative tasks. We model the data with flexible additive energy distributions, where each energy term represents an attribute, and derive a simple alternative to empirical risk minimization termed compositional risk minimization (CRM). We first train an additive energy classifier to predict the multiple attributes and then adjust this classifier to tackle compositional shifts. We provide an extensive theoretical analysis of CRM, where we show that our proposal extrapolates to special affine hulls of seen attribute combinations. Empirical evaluations on benchmark datasets confirms the improved robustness of CRM compared to other methods from the literature designed to tackle various forms of subpopulation shifts.

组合泛化分布偏移能量模型鲁棒学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。