测试大模型在不同表述下是否能灵活切换立场表达,揭示其可信对话能力。
Epistemic Stance Flexibility Probing: Measuring Prompt-Conditioned Register Shift in Large Language Models

- 设计对比实验,让模型区分外部引用和自我表态两种提问方式。
- 8个主流模型表现差异大,大模型不等于立场灵活,推理优化也未必提升。
- 重点看立场密度而非表面词,适合评估模型可信度与安全对话能力。
语言模型可能被问及专家对争议性问题的看法,或它自身对该问题的立场。一个可信的对话代理应能区分这两类请求,并分别采用中立引用和立场表达的语用风格。现有评测基准无法直接衡量这种立场转换能力。我们提出ESFP,一种以外部引用与自我归因提示为基本单位的行为评测框架,包含104个精心控制的题目,覆盖六个认知范畴和五种句式模板,从四个维度评估模型响应:词汇级自我归属、层级响应角色框架、由大模型评审团评估的句子级立场内容密度,以及跨条件立场一致性。评估八个来自五个厂商的前沿模型发现,认知灵活性与通用能力基本无关:270亿参数的开源模型媲美最强专有系统,某家族旗舰模型反而弱于轻量版,推理优化模型也未表现出更高灵活性。立场内容密度是最强信号,而表面词如‘我认为’变化大但立场未必改变。我们提供逐项置信区间、权重敏感性分析及综合评分解释局限性的讨论。ESFP衡量的是模型在不同归因条件下调整认知立场的倾向性,而非通用能力。
原文摘要 · Abstract (English)
A language model may be asked either what experts believe about a contested claim or what it believes about the claim itself. A trustworthy conversational agent should distinguish these two requests and respond in different epistemic registers: neutral attribution in the first case and stance expression in the second. Whether such a shift occurs-and whether it occurs coherently-is not directly assessed by existing benchmarks for accuracy, instruction following, or safety. We introduce ESFP, a behavioral benchmark that treats the contrast between externally attributed and self-attributed prompts as the fundamental unit of measurement. ESFP consists of 104 carefully controlled items spanning six epistemic categories and five phrasing templates, and evaluates model responses along four complementary dimensions: lexical self-attribution, representation-level responsiveness to role framing, sentence-level stance content density assessed by an LLM judge panel, and cross-condition stance consistency. Evaluating eight frontier models from five vendors, we find that epistemic flexibility is largely orthogonal to general model capability: a 27B open-weight model matches the strongest proprietary systems, the flagship model of one family underperforms its lightweight counterpart, and reasoning-optimized models do not consistently exhibit higher flexibility. Stance content density provides the strongest signal, while surface-level lexical markers such as 'I think' can change substantially without corresponding changes in expressed stance. We provide item-level bootstrap confidence intervals, weight-sensitivity analyses, and an explicit discussion of the interpretation limits of the composite score. ESFP measures a model's propensity to adapt its epistemic stance under changing attribution conditions, rather than a general competence measure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。