arXiv:2601.13537cs.CLcs.AI2026-01被引 9

提示词措辞能显著影响大模型评判结果,需警惕评估中的框架偏差。

When Wording Steers the Evaluation: Framing Bias in LLM judges

  • 用正反谓语结构设计对称提示,测试模型评判受措辞影响程度。
  • 14个大模型在高风险任务中均表现出明显框架依赖性。
  • 研究揭示评估系统存在结构性偏差,适合关注AI评测可靠性的研究者阅读。

大语言模型(LLMs)的响应常因提示词措辞而变化,表明细微的表述引导可改变其回答。然而,这种框架偏差对基于大模型的评估(要求判断稳定且公正)的影响尚未充分探索。受心理学中框架效应启发,我们系统研究了故意的提示框架如何在四个高风险评估任务中扭曲模型判断。通过使用谓语正向和负向结构设计对称提示,我们证明此类框架会引发模型输出的显著差异。在14个大模型裁判中,均观察到明显的框架敏感性,不同模型家族表现出不同的倾向性:有的更倾向于同意,有的则更倾向于拒绝。这些发现表明,框架偏差是当前基于大模型评估系统的一种结构性特征,凸显了采用框架感知协议的必要性。

原文摘要 · Abstract (English)

Large language models (LLMs) are known to produce varying responses depending on prompt phrasing, indicating that subtle guidance in phrasing can steer their answers. However, the impact of this framing bias on LLM-based evaluation, where models are expected to make stable and impartial judgments, remains largely underexplored. Drawing inspiration from the framing effect in psychology, we systematically investigate how deliberate prompt framing skews model judgments across four high-stakes evaluation tasks. We design symmetric prompts using predicate-positive and predicate-negative constructions and demonstrate that such framing induces significant discrepancies in model outputs. Across 14 LLM judges, we observe clear susceptibility to framing, with model families showing distinct tendencies toward agreement or rejection. These findings suggest that framing bias is a structural property of current LLM-based evaluation systems, underscoring the need for framing-aware protocols.

大模型评估框架偏差提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。