精心挑选例子能让GPT-3.5在作文评分中超越部分GPT-4模型。
The Impact of Example Selection in Few-Shot Prompting on Automated Essay Scoring Using GPT Models
- 通过对比不同例子组合与顺序,测试GPT在少样本提示下的评分表现。
- GPT-3.5受例子影响更大,存在多数标签和最近例子偏差,而GPT-4更稳定。
- 2023年6月版GPT-4表现最优,说明小版本也需单独评估。
本研究探讨了在使用GPT模型进行少样本提示时,示例选择对自动作文评分(AES)性能的影响。实验涉及119个包含不同示例的提示,采用二次加权克劳氏一致性系数(QWK)衡量GPT与人工评分者的一致性。回归分析用于量化示例选择带来的偏差。结果表明,示例选择对GPT-3.5的影响大于GPT-4,且两者均表现出多数标签偏差和最近示例偏差,其中GPT-3.5更为显著。值得注意的是,合理选择示例可使GPT-3.5表现超越部分GPT-4模型。在各版本中,2023年6月的GPT-4展现出最高的稳定性与性能。研究揭示了在少样本提示下示例选择的重要性,尤其对GPT-3.5,强调需对每个模型版本独立评估。
原文摘要 · Abstract (English)
This study investigates the impact of example selection on the performance of au-tomated essay scoring (AES) using few-shot prompting with GPT models. We evaluate the effects of the choice and order of examples in few-shot prompting on several versions of GPT-3.5 and GPT-4 models. Our experiments involve 119 prompts with different examples, and we calculate the quadratic weighted kappa (QWK) to measure the agreement between GPT and human rater scores. Regres-sion analysis is used to quantitatively assess biases introduced by example selec-tion. The results show that the impact of example selection on QWK varies across models, with GPT-3.5 being more influenced by examples than GPT-4. We also find evidence of majority label bias, which is a tendency to favor the majority la-bel among the examples, and recency bias, which is a tendency to favor the label of the most recent example, in GPT-generated essay scores and QWK, with these biases being more pronounced in GPT-3.5. Notably, careful example selection enables GPT-3.5 models to outperform some GPT-4 models. However, among the GPT models, the June 2023 version of GPT-4, which is not the latest model, exhibits the highest stability and performance. Our findings provide insights into the importance of example selection in few-shot prompting for AES, especially in GPT-3.5 models, and highlight the need for individual performance evaluations of each model, even for minor versions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。