对比BERT与GPT在政治科学文本分类中的表现,指导小数据场景下模型选择。
Selecting Between BERT and GPT for Text Classification in Political Science Research
- 用提示工程+GPT做少样本学习,探索早期研究可行性。
- 当数据量达1000样本时,BERT微调性能显著优于GPT零/少样本。
- 适合资源有限、标签数据少的政治学定量文本研究者参考。
政治科学家常面临文本分类中的数据稀缺问题。近年来,微调BERT及其变体被广泛认为是有效解决方案。本文探讨结合提示工程的GPT模型作为替代方案的潜力。我们在不同类别数和复杂度的任务上进行系列实验,评估低数据场景下BERT与GPT模型的分类效果。结果表明,尽管GPT在零样本和少样本学习中表现合理,适合初期探索,但其整体性能通常不及或仅匹配BERT微调,尤其当训练集达到1000样本以上时差距明显。我们进一步从性能、易用性和成本角度比较两种方法,为面临数据限制的研究者提供实用建议。研究结果对低资源环境下开展定量文本分析具有重要参考价值。
原文摘要 · Abstract (English)
Political scientists often grapple with data scarcity in text classification. Recently, fine-tuned BERT models and their variants have gained traction as effective solutions to address this issue. In this study, we investigate the potential of GPT-based models combined with prompt engineering as a viable alternative. We conduct a series of experiments across various classification tasks, differing in the number of classes and complexity, to evaluate the effectiveness of BERT-based versus GPT-based models in low-data scenarios. Our findings indicate that while zero-shot and few-shot learning with GPT models provide reasonable performance and are well-suited for early-stage research exploration, they generally fall short - or, at best, match - the performance of BERT fine-tuning, particularly as the training set reaches a substantial size (e.g., 1,000 samples). We conclude by comparing these approaches in terms of performance, ease of use, and cost, providing practical guidance for researchers facing data limitations. Our results are particularly relevant for those engaged in quantitative text analysis in low-resource settings or with limited labeled data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。