arXiv:2601.03450cs.LGcs.AI2026-01

让模型轻松理解用户自定义分类任务,零样本也能准确识别新类别。

Soft Contextualized Encoder For User Defined Text Classification

  • 用软提示和标签集上下文增强输入编码,提升对新类别的理解能力。
  • 在多个未见数据集上表现领先,零样本分类准确率显著优于基线。
  • 适合企业分析、内容审核等需要快速适应新分类场景的应用。

用户自定义文本分类(UDTC)旨在将输入文本分类到用户指定的、此前未见过的类别中,这一场景在企业分析、内容审核和领域信息检索中频繁出现。本文提出一种软上下文编码器架构用于UDTC,通过将每个候选标签与标签集及输入查询的静态软提示表示进行上下文关联,实现更精准的语义对齐。在多源、多样数据集上训练后,模型能有效泛化至任意领域中完全未见的话题集合,实现零样本分类。我们在保留的分布内测试数据及多个未见的UDTC基准上评估该架构,结果表明,该模型在各数据集上均达到当前最优性能,持续超越或匹配现有基线方法。

原文摘要 · Abstract (English)

User-Defined Text Classification (UDTC) considers the challenge of classifying input text to user-specified, previously unseen classes, a setting that arises frequently in real-world applications such as enterprise analytics, content moderation, and domain-specific information retrieval. We propose a soft-contextualized encoder architecture for UDTC which contextualizes each candidate label with the label set and a static soft prompt representation of the input query. Training on diverse, multi-source datasets enables the model to generalize effectively to zero-shot classification over entirely unseen topic sets drawn from arbitrary domains. We evaluate the proposed architecture both on held-out in-distribution test data and on multiple unseen UDTC benchmarks. Across datasets, the model achieves state-of-the-art performance, consistently outperforming or matching the baselines.

文本分类零样本学习用户自定义

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。