首个多语言文本分类开放集学习与发现基准,支持12种语言的未知类别识别。
MOSLD-Bench: Multilingual Open-Set Learning and Discovery Benchmark for Text Categorization
- 构建多语言数据集,整合现有与新采集新闻数据,覆盖12语言共96万样本。
- 提出分阶段持续发现与学习框架,实现测试时未知类别的自动识别与归类。
- 提供可复现基准,适合研究开放世界文本分类与跨语言模型泛化任务者使用。
开放集学习与发现(OSLD)是机器学习中一项挑战性任务,即在测试阶段可能出现未知类别的样本。它可视为零样本学习的推广,其中新类别事先未知,需主动发现。尽管零样本学习在文本分类中已有广泛研究,尤其得益于预训练语言模型的发展,但开放集学习与发现在文本领域仍属新兴方向。为此,我们提出了首个多语言开放集学习与发现(MOSLD)基准,用于主题分类任务,涵盖12种语言共96万条数据样本。构建过程中,我们重新组织现有数据集并从新闻领域收集新样本。此外,我们提出一种新型多阶段融合框架,用于持续发现并学习新类别。评估了多种语言模型(包括自研模型)以提供未来研究参考。基准代码已开源:https://github.com/Adriana19Valentina/MOSLD-Bench。
原文摘要 · Abstract (English)
Open-set learning and discovery (OSLD) is a challenging machine learning task in which samples from new (unknown) classes can appear at test time. It can be seen as a generalization of zero-shot learning, where the new classes are not known a priori, hence involving the active discovery of new classes. While zero-shot learning has been extensively studied in text classification, especially with the emergence of pre-trained language models, open-set learning and discovery is a comparatively new setup for the text domain. To this end, we introduce the first multilingual open-set learning and discovery (MOSLD) benchmark for text categorization by topic, comprising 960K data samples across 12 languages. To construct the benchmark, we (i) rearrange existing datasets and (ii) collect new data samples from the news domain. Moreover, we propose a novel framework for the OSLD task, which integrates multiple stages to continuously discover and learn new classes. We evaluate several language models, including our own, to obtain results that can be used as reference for future work. We release our benchmark at https://github.com/Adriana19Valentina/MOSLD-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。