arXiv:2609.07448cs.CL2026-09

测试大模型如何被问题表述方式影响,发现同一事实下答案可能因提问方式不同而变化。

FramingQA: Does the Question Shape the Answer? Measuring the Compositional Framing Effect

论文配图:FramingQA: Does the Question Shape the Answer? Measuring the Compositional Framing Effect
图 1 · 摘自论文原文
  • 设计三重嵌套框架测试模型对问题表述的敏感性
  • 9个模型在相同事实下因提问方式不同,答案差异显著
  • 适合关注AI决策公平性与可靠性的研究人员

我们提出FramingQA基准,用于衡量大语言模型在法律、医疗、金融和机器人模拟领域对问题表述的敏感性。大模型常因用户隐含立场的细微措辞变化而改变回答,导致建议受提问方式而非事实影响,在高风险领域后果严重。在真实场景中,专家与非专家常提出包含不完整或误导性假设的问题,使模型极易受表述偏见影响。为此,我们在三个嵌套层级注入表述偏见:问题本身(根)、附加至中立问题的偏见前提(命题)、以及偏见前提与偏见问题配对(全局)。评估了四个系列共九个开源模型(3.8B-70B),发现单一变体上的高准确率并不能保证在不同表述下保持一致输出。

原文摘要 · Abstract (English)

We introduce FramingQA, a benchmark that measures the model sensitivity to question framing across law, medicine, finance, and robotic simulations. Large language models (LLMs) often change their responses to subtle rephrasings that align with an implied stance by users. This can leave users with advice tainted by how they happened to phrase a question rather than by the underlying facts, and the consequences are highly costly in high-stakes domains. Because in the realistic scenarios, both expert practitioners and non-expert users frequently ask LLMs questions containing incomplete or misleading assumptions, models are highly susceptible to those framings. To test this, we inject the framing bias across three nested levels: a framing-biased question phrasing (root), an injected framing-biased premise prepended to a neutral question (propositional), and a premise paired with a framing-biased question (global). Evaluating nine open models (3.8B-70B) across four families, we find that strong per-variant accuracy does not guarantee the robustness across differently phrased questions under the fixed factual information.

大模型评估表述偏见语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。