对比人类与大模型对提示词变化的敏感度,发现两者均有脆性但表现不同。
Are Humans as Brittle as Large Language Models?
- 用相同指令测试人类与大模型在文本分类任务中的表现差异。
- 替换标签集或格式时,人与模型均更易出错,但模型更敏感于拼写错误和标签顺序。
- 揭示大模型提示脆性可能反映真实人类标注变异性,非纯粹缺陷。
大语言模型(LLMs)的输出不稳定,源于解码过程的非确定性及提示词脆性。尽管模型生成的不确定性可能通过输出分布偏移模拟人类标注中的固有不确定性,但提示词脆性是否仅属于模型尚无定论。本研究系统比较了提示修改对大模型与人类标注者的影响,聚焦于人类是否同样敏感于提示扰动。在一系列文本分类任务中,对人类与大模型施加相同的提示变化。结果表明,当标签集或标签格式被替换时,人类与大模型均表现出更强的脆性;然而,人类判断受拼写错误和标签顺序颠倒的影响较小,而大模型则更为敏感。这提示提示词脆性可能部分反映了真实的人类标注变异,而非单纯的问题。
原文摘要 · Abstract (English)
The output of large language models (LLMs) is unstable, due both to non-determinism of the decoding process as well as to prompt brittleness. While the intrinsic non-determinism of LLM generation may mimic existing uncertainty in human annotations through distributional shifts in outputs, it is largely assumed, yet unexplored, that the prompt brittleness effect is unique to LLMs. This raises the question: do human annotators show similar sensitivity to prompt changes? If so, should prompt brittleness in LLMs be considered problematic? One may alternatively hypothesize that prompt brittleness correctly reflects human annotation variances. To fill this research gap, we systematically compare the effects of prompt modifications on LLMs and identical instruction modifications for human annotators, focusing on the question of whether humans are similarly sensitive to prompt perturbations. To study this, we prompt both humans and LLMs for a set of text classification tasks conditioned on prompt variations. Our findings indicate that both humans and LLMs exhibit increased brittleness in response to specific types of prompt modifications, particularly those involving the substitution of alternative label sets or label formats. However, the distribution of human judgments is less affected by typographical errors and reversed label order than that of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。