LLM评价结果受来源表述影响,中国作者声明导致评分显著下降
Source framing triggers systematic evaluation bias in Large Language Models
- 通过改变文本来源归属(人类或LLM),测试模型评价一致性
- 中国作者标签使所有模型评分降低,深求推理机受影响最明显
- 揭示评估系统存在隐蔽偏见,适用于审查AI决策公平性研究
大型语言模型(LLMs)不仅用于生成文本,也越来越多地用于文本评估,这引发了对其判断是否一致、无偏且对框架效应具有鲁棒性的紧迫疑问。本研究系统考察了四种前沿模型(OpenAI o3-mini、Deepseek Reasoner、xAI Grok 2 和 Mistral)在24个社会、政治及公共卫生相关主题上对4,800条叙事陈述的评估表现,共完成192,000次评估。我们通过操纵每条陈述的披露来源,评估将其归因于另一大模型或特定国籍的人类作者如何影响评估结果。在盲测条件下,不同模型在各主题上表现出高度的跨模型与模型内一致性。然而,引入来源框架后,这种一致性被打破:将陈述归因于中国个体显著降低了所有模型的评估得分,尤其影响深求推理机(Deepseek Reasoner)。研究揭示,框架效应会深刻影响文本评估,对基于大模型的信息系统的完整性、中立性与公平性具有重要影响。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used not only to generate text but also to evaluate it, raising urgent questions about whether their judgments are consistent, unbiased, and robust to framing effects. In this study, we systematically examine inter- and intra-model agreement across four state-of-the-art LLMs (OpenAI o3-mini, Deepseek Reasoner, xAI Grok 2, and Mistral) tasked with evaluating 4,800 narrative statements on 24 different topics of social, political, and public health relevance, for a total of 192,000 assessments. We manipulate the disclosed source of each statement to assess how attribution to either another LLM or a human author of specified nationality affects evaluation outcomes. We find that, in the blind condition, different LLMs display a remarkably high degree of inter- and intra-model agreement across topics. However, this alignment breaks down when source framing is introduced. Here we show that attributing statements to Chinese individuals systematically lowers agreement scores across all models, and in particular for Deepseek Reasoner. Our findings reveal that framing effects can deeply affect text evaluation, with significant implications for the integrity, neutrality, and fairness of LLM-mediated information systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。