用小样本训练出高效开源的政治文本分类模型,性能超大模型。
Political DEBATE: Efficient Zero-shot and Few-shot Classifiers for Political Text
- 基于随机抽样10-25篇文档微调DeBERTa,实现高效零/少样本分类。
- 仅用少量数据训练的模型,超越千份数据训练的监督模型。
- 模型开源且高效,适合追求可复现性的政治文本研究者使用。
社会科学家迅速采用大语言模型进行无监督文档标注,即零样本学习。但其高计算成本、高昂费用及专有性质常与可复现性及开放科学标准相悖。本文提出Political DEBATE(DeBERTa用于文本蕴含的算法)模型,用于政治文本的零样本与少样本分类。该模型在零/少样本分类任务中表现不逊于甚至优于当前最优的大语言模型,且效率高出数个数量级,完全开源。仅需对10-25篇文档进行简单随机采样训练,即可超越使用数百至数千份数据训练的监督分类器及使用复杂工程提示的生成式模型。此外,我们发布了用于训练的PolNLI数据集——一个包含超过20万条政治文档、覆盖800余项分类任务的高质量标注语料库。
原文摘要 · Abstract (English)
Social scientists quickly adopted large language models due to their ability to annotate documents without supervised training, an ability known as zero-shot learning. However, due to their compute demands, cost, and often proprietary nature, these models are often at odds with replication and open science standards. This paper introduces the Political DEBATE (DeBERTa Algorithm for Textual Entailment) language models for zero-shot and few-shot classification of political documents. These models are not only as good, or better than, state-of-the art large language models at zero and few-shot classification, but are orders of magnitude more efficient and completely open source. By training the models on a simple random sample of 10-25 documents, they can outperform supervised classifiers trained on hundreds or thousands of documents and state-of-the-art generative models with complex, engineered prompts. Additionally, we release the PolNLI dataset used to train these models -- a corpus of over 200,000 political documents with highly accurate labels across over 800 classification tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。