arXiv:2502.12965cs.CLcs.AI2025-02Conference of the …综述被引 2

文本分类中类别分布随时间变化,这篇综述系统梳理了应对方法。

Text Classification Under Class Distribution Shift: A Survey

  • 按分布偏移类型分三类:Universum学习、零样本学习、开放集学习
  • 总结主流缓解策略,涵盖不同设定下的关键方法
  • 指出持续学习是解决分布漂移的核心方向,适合关注模型泛化者

机器学习模型通常假设训练与测试数据来自同一分布,但实际应用中这一假设常被打破——测试数据的分布会随时间变化,制约传统模型应用。文本分类领域尤为明显,因新话题不断涌现。本文综述研究开放集文本分类及相关任务的文献,依据分布偏移的约束条件和对应问题建模方式,将方法分为学习与Universum、零样本学习、开放集学习三类。接着讨论各类设定下的主要缓解策略。进一步提出若干未来研究方向,旨在突破现有技术水平。最后说明持续学习如何有效应对类别分布变化带来的挑战。相关论文列表见 https://github.com/Eduard6421/Open-Set-Survey。

原文摘要 · Abstract (English)

The basic underlying assumption of machine learning (ML) models is that the training and test data are sampled from the same distribution. However, in daily practice, this assumption is often broken, i.e. the distribution of the test data changes over time, which hinders the application of conventional ML models. One domain where the distribution shift naturally occurs is text classification, since people always find new topics to discuss. To this end, we survey research articles studying open-set text classification and related tasks. We divide the methods in this area based on the constraints that define the kind of distribution shift and the corresponding problem formulation, i.e. learning with the Universum, zero-shot learning, and open-set learning. We next discuss the predominant mitigation approaches for each problem setup. We further identify several future work directions, aiming to push the boundaries beyond the state of the art. Finally, we explain how continual learning can solve many of the issues caused by the shifting class distribution. We maintain a list of relevant papers at https://github.com/Eduard6421/Open-Set-Survey.

文本分类分布偏移开放集学习持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。