arXiv:2411.18924cs.CL2024-11中稿 · AIED2024被引 18

精心挑选例子能让GPT-3.5在作文评分中超越部分GPT-4模型。

The Impact of Example Selection in Few-Shot Prompting on Automated Essay Scoring Using GPT Models

  • 通过对比不同例子组合与顺序,测试GPT在少样本提示下的评分表现。
  • GPT-3.5受例子影响更大,存在多数标签和最近例子偏差,而GPT-4更稳定。
  • 2023年6月版GPT-4表现最优,说明小版本也需单独评估。

本研究探讨了在使用GPT模型进行少样本提示时,示例选择对自动作文评分(AES)性能的影响。实验涉及119个包含不同示例的提示,采用二次加权克劳氏一致性系数(QWK)衡量GPT与人工评分者的一致性。回归分析用于量化示例选择带来的偏差。结果表明,示例选择对GPT-3.5的影响大于GPT-4,且两者均表现出多数标签偏差和最近示例偏差,其中GPT-3.5更为显著。值得注意的是,合理选择示例可使GPT-3.5表现超越部分GPT-4模型。在各版本中,2023年6月的GPT-4展现出最高的稳定性与性能。研究揭示了在少样本提示下示例选择的重要性,尤其对GPT-3.5,强调需对每个模型版本独立评估。

原文摘要 · Abstract (English)

This study investigates the impact of example selection on the performance of au-tomated essay scoring (AES) using few-shot prompting with GPT models. We evaluate the effects of the choice and order of examples in few-shot prompting on several versions of GPT-3.5 and GPT-4 models. Our experiments involve 119 prompts with different examples, and we calculate the quadratic weighted kappa (QWK) to measure the agreement between GPT and human rater scores. Regres-sion analysis is used to quantitatively assess biases introduced by example selec-tion. The results show that the impact of example selection on QWK varies across models, with GPT-3.5 being more influenced by examples than GPT-4. We also find evidence of majority label bias, which is a tendency to favor the majority la-bel among the examples, and recency bias, which is a tendency to favor the label of the most recent example, in GPT-generated essay scores and QWK, with these biases being more pronounced in GPT-3.5. Notably, careful example selection enables GPT-3.5 models to outperform some GPT-4 models. However, among the GPT models, the June 2023 version of GPT-4, which is not the latest model, exhibits the highest stability and performance. Our findings provide insights into the importance of example selection in few-shot prompting for AES, especially in GPT-3.5 models, and highlight the need for individual performance evaluations of each model, even for minor versions.

自动评分少样本学习GPT模型提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。