arXiv:2510.06997cs.CYcs.AI2025-10中稿 · T4E 2025 for poste…

测试发现,详细提示无法让AI评分更稳定,挑战了其像人一样评估的假设。

The Limits of Goal-Setting Theory in LLM-Driven Assessment

  • 用4种逐步细化的提示让ChatGPT评估29份学生作业,测试提示具体性对评分一致性的影响。
  • 尽管提示越来越具体,评分一致性(Cohen's Kappa)未显著提升,波动依然明显。
  • 研究揭示当前大模型在评估任务中缺乏人类评卷的可预测性,适合关注AI评估可靠性的研究者。

许多用户在使用ChatGPT等AI工具时,会采用一种将系统视为类人实体的心理模型,我们称之为Model H。根据目标设定理论,目标越具体,表现差异应越小。若Model H成立,则给聊天机器人提供更详细的指令应带来更一致的评价行为。本文通过一项控制实验,在该实验中,ChatGPT使用四种提示度逐步增加的指令评估29份学生提交作业,以重复运行中的评分者内一致性(Cohen's Kappa)衡量一致性。结果与预期相反:随着提示具体性提升,性能并未持续改善,表现方差基本保持不变。这一发现质疑了大模型行为类似于人类评估者的假设,并凸显未来模型开发中需增强鲁棒性及输入整合能力。

原文摘要 · Abstract (English)

Many users interact with AI tools like ChatGPT using a mental model that treats the system as human-like, which we call Model H. According to goal-setting theory, increased specificity in goals should reduce performance variance. If Model H holds, then prompting a chatbot with more detailed instructions should lead to more consistent evaluation behavior. This paper tests that assumption through a controlled experiment in which ChatGPT evaluated 29 student submissions using four prompts with increasing specificity. We measured consistency using intra-rater reliability (Cohen's Kappa) across repeated runs. Contrary to expectations, performance did not improve consistently with increased prompt specificity, and performance variance remained largely unchanged. These findings challenge the assumption that LLMs behave like human evaluators and highlight the need for greater robustness and improved input integration in future model development.

AI评估大模型提示工程可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。