通过优化提示词结构,显著提升大模型在社科文本分类中的准确率。
Navigating the Prompt Space: Improving LLM Classification of Social Science Texts Through Prompt Engineering
- 系统调整标签描述、指令引导和少样本示例来优化提示词
- 适度增加提示上下文可大幅提升性能,过度增加反而可能降低准确率
- 不同模型、任务和批量大小表现差异大,需单独验证
近年来,使用大语言模型(LLMs)进行社会科学文本分类的研究表明,可大幅降低计算成本,性能有时甚至媲美传统方法。然而,现有测试中性能波动较大,因此我们关注如何最大化分类效果。本文聚焦提示词上下文作为提升准确率的潜在途径,系统性地调整三个提示工程维度:标签描述、指令引导和少样本示例。在两个不同任务上的实验显示,适度增加提示上下文能带来最大性能提升,进一步增加仅产生边际收益;令人担忧的是,过度增加上下文有时反而会降低准确率。此外,实验还揭示了模型、任务和批次大小之间的显著异质性,强调每项大模型编码任务都需独立验证,不可依赖通用规则。
原文摘要 · Abstract (English)
Recent developments in text classification using Large Language Models (LLMs) in the social sciences suggest that costs can be cut significantly, while performance can sometimes rival existing computational methods. However, with a wide variance in performance in current tests, we move to the question of how to maximize performance. In this paper, we focus on prompt context as a possible avenue for increasing accuracy by systematically varying three aspects of prompt engineering: label descriptions, instructional nudges, and few shot examples. Across two different examples, our tests illustrate that a minimal increase in prompt context yields the highest increase in performance, while further increases in context only tend to yield marginal performance increases thereafter. Alarmingly, increasing prompt context sometimes decreases accuracy. Furthermore, our tests suggest substantial heterogeneity across models, tasks, and batch size, underlining the need for individual validation of each LLM coding task rather than reliance on general rules.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。