arXiv:2607.10202cs.LG2026-07

主流评测方法会掩盖模型位置偏见,导致中立性误判。

When Counterbalancing Hides the Bias: Access-Conditioned Position Lock in Forced-Choice LLM Evaluation

  • 用反向平衡法测量模型态度时,位置锁定会被误判为中立
  • 9个模型中位置锁定比例与评估指数高度相关(r=-0.986)
  • 需附加诊断工具,否则易误判模型无立场

强制选择探测采用反向平衡设计,是衡量语言模型“价值取向”的标准方法,通过重复测试的集中度/极端度指数来判断模型承诺程度。我们发现该估计算法在低值端不可识别:反向平衡本意消除位置偏见,却将位置锁定(模型无论内容均选同一选项)映射为接近0.5的中立信号,使“软弱”与“非参与”无法区分。在九个模型中,位置锁定项占比与指数几乎完美负相关(r = -0.986),这是结构性结果而非偶然发现。真正有意义的是突破此边界后的剩余承诺。三个被指数标记为“软”的模型实为最严重的位置锁定;当访问配置允许推理(部署客户端→原始API→启用推理)时,指数从0.06升至0.63、0.34升至0.66再至0.64,而锁定消失(Opus:0.61→0.39→0.22),证明低读数是可识别性失败,非真实态度。启用推理的读数并非“真实”价值;关键在于指数本身无法在低值端识别承诺。访问配置(部署客户端、推理开关)是该失效的重要来源,已在两种路径(Anthropic订阅CLI与DeepSeek客户端)中验证,且在冻结基准中与提供方混杂。我们提出必须伴随位置锁定诊断,否则浓度盲审可能误报中立;方向翻转成分虽受干扰,但仍能可靠识别跨模型真实分歧。

原文摘要 · Abstract (English)

Forced-choice probes with counterbalanced orientations are a standard tool for measuring language-model "value dispositions," and a concentration/extremity index over repeated draws is read as how sharply a model commits. We show this estimator is not identifiable at its low end: counterbalancing, meant to remove position bias, instead maps a position-lock (a model returning the same answer letter regardless of content) onto the same near-0.5 signature as genuine neutrality, so "softness" and "non-engagement" cannot be distinguished from a content-independent letter-bias. Across nine models the fraction of position-locked items tracks the index almost perfectly (r=-0.986) - a structural consequence, not a finding: the informative quantity is the residual from that bound, the commitment a model shows on the items it does engage. The three models the index reads as soft are the most position-locked, and the lock resolves under access configurations that permit reasoning (as-deployed CLI -> raw API -> reasoning-enabled), moving the index 0.06->0.63 and 0.34->0.66->0.64 while lock collapses (Opus: 0.61->0.39->0.22); this establishes the low readings as an identifiability failure, not a disposition. The reasoning-enabled reading is not a "true" value either; the point is that the index alone cannot identify commitment at its low end. Access configuration (deployment client, reasoning on/off) is one generator of this failure, shown on two access paths (an Anthropic subscription CLI and a DeepSeek client), and in a frozen benchmark it is confounded with provider. We contribute a position-lock diagnostic that must accompany any concentration reading, and show that a concentration-blind audit risks reporting neutrality where a reasoning-permitting condition yields concentrated choices. The direction-flip component, largely robust to the artifact, still identifies genuine cross-model disagreement.

模型评测位置偏见注意力机制语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。