arXiv:2502.17091cs.CLcs.AI2025-02被引 5

对比人类与大模型在真实文本中的框架效应,发现两者反应相似但大模型差异显著。

Comparing the Framing Effect in Humans and LLMs on Naturally Occurring Texts

  • 构建真实文本数据集WildFrame,对比人类与大模型对正负框架的响应
  • 11个大模型均呈现人类类似行为(相关性≥0.52),且更受正向框架影响
  • GPT系列与人类相关性最低,引发对模型应否模仿人类认知偏见的讨论

人类会受信息呈现方式影响,这种现象称为框架效应。已有研究认为大模型也可能受此影响,但依赖合成数据且未与人类行为对比。为此,我们提出WildFrame——一个用于评估大模型在自然语料中对正负框架响应的数据集,同时收集人类对相同数据的标注。WildFrame包含1,000条真实文本,每条均表达明确情感,再以正/负视角重述,并获取人类情感判断。在该数据集上评估11个大模型,发现所有模型均表现出类人反应(相关性≥0.52),且人类与模型均更易受正向框架影响。值得注意的是,GPT系列模型与人类行为相关性最低。这一结果引发关于当前大模型发展目标的讨论:是否应贴近人类认知偏差以保留如框架效应等心理现象,还是应消除此类偏见以实现公平与一致。

原文摘要 · Abstract (English)

Humans are influenced by how information is presented, a phenomenon known as the framing effect. Prior work suggests that LLMs may also be susceptible to framing, but it has relied on synthetic data and did not compare to human behavior. To address this gap, we introduce WildFrame - a dataset for evaluating LLM responses to positive and negative framing in naturally-occurring sentences, alongside human responses on the same data. WildFrame consists of 1,000 real-world texts selected to convey a clear sentiment; we then reframe each text in either a positive or negative light and collect human sentiment annotations. Evaluating eleven LLMs on WildFrame, we find that all models respond to reframing in a human-like manner ($r\geq0.52$), and that both humans and models are influenced more by positive than negative reframing. Notably, GPT models are the least correlated with human behavior among all tested models. These findings raise a discussion around the goals of state-of-the-art LLM development and whether models should align closely with human behavior, to preserve cognitive phenomena such as the framing effect, or instead mitigate such biases in favor of fairness and consistency.

框架效应大模型行为人类对比认知偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。