arXiv:2502.11830cs.CL2025-02被引 25

对比大模型在多语言文本分类中的表现,发现合成数据更优,零样本适用有限。

Text Classification in the LLM Era -- Where do we stand?

  • 用32个跨8语种数据集,比较零样本、小模型微调与合成数据方法
  • 合成数据来自多个大模型,性能优于零样本开放大模型
  • 跨语言差异大,指导多语言分类系统设计

大型语言模型(LLM)革新了自然语言处理,在多项任务中表现出显著提升。本文研究了这类模型在文本分类中的作用,并与依赖小型预训练模型的方法进行对比。基于涵盖8种语言的32个数据集,我们比较了零样本分类、少样本微调和基于合成数据的分类器与使用完整人工标注数据构建的分类器。结果显示,零样本方法在情感分类中表现良好,但在其他任务中被超越;而来自多个LLM的合成数据可构建出比零样本开放大模型更好的分类器。此外,所有分类场景中均存在显著的跨语言性能差异。这些发现有助于指导跨语言文本分类系统的开发。

原文摘要 · Abstract (English)

Large Language Models revolutionized NLP and showed dramatic performance improvements across several tasks. In this paper, we investigated the role of such language models in text classification and how they compare with other approaches relying on smaller pre-trained language models. Considering 32 datasets spanning 8 languages, we compared zero-shot classification, few-shot fine-tuning and synthetic data based classifiers with classifiers built using the complete human labeled dataset. Our results show that zero-shot approaches do well for sentiment classification, but are outperformed by other approaches for the rest of the tasks, and synthetic data sourced from multiple LLMs can build better classifiers than zero-shot open LLMs. We also see wide performance disparities across languages in all the classification scenarios. We expect that these findings would guide practitioners working on developing text classification systems across languages.

文本分类大模型多语言合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。