arXiv:2512.01152cs.LGcs.AI2025-12

提出CoLOR方法,在背景分布变化时仍能可靠识别新类别。

Open-Set Domain Adaptation Under Background Distribution Shift: Challenges and A Provably Efficient Solution

  • 基于类别可分性假设,设计理论保证的开放集识别新方法
  • 在过参数化简化场景下证明性能优于基线方法
  • 首次系统分析新类规模对性能的影响,适合实际部署场景

当机器学习系统部署于真实世界时,需应对数据分布随时间变化的挑战。此类变化包括训练中未出现的新类别(即开放集识别),以及已知类别的分布改变。现有开放集识别的理论保证大多基于已知类别分布(称作背景分布)固定的假设。本文提出CoLOR方法,在背景分布发生变化的困难情形下仍能保证解决开放集识别问题。我们在新类别与非新类别可分的合理假设下,给出了理论保证,并在简化过参数化设置下证明其性能优于代表性基线。通过开发可扩展且鲁棒的技术,我们在图像和文本数据上进行了全面实证评估。结果表明,CoLOR在背景分布转移下显著优于现有方法。此外,我们还揭示了新类规模对性能的影响机制,这一因素在以往研究中尚未被充分探讨。

原文摘要 · Abstract (English)

As we deploy machine learning systems in the real world, a core challenge is to maintain a model that is performant even as the data shifts. Such shifts can take many forms: new classes may emerge that were absent during training, a problem known as open-set recognition, and the distribution of known categories may change. Guarantees on open-set recognition are mostly derived under the assumption that the distribution of known classes, which we call the background distribution, is fixed. In this paper we develop CoLOR, a method that is guaranteed to solve open-set recognition even in the challenging case where the background distribution shifts. We prove that the method works under benign assumptions that the novel class is separable from the non-novel classes, and provide theoretical guarantees that it outperforms a representative baseline in a simplified overparameterized setting. We develop techniques to make CoLOR scalable and robust, and perform comprehensive empirical evaluations on image and text data. The results show that CoLOR significantly outperforms existing open-set recognition methods under background shift. Moreover, we provide new insights into how factors such as the size of the novel class influences performance, an aspect that has not been extensively explored in prior work.

开放集识别领域自适应分布偏移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。