用量子问卷等式审计大模型问答顺序效应,发现测量饱和会扭曲结果
Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat
- 基于量子问卷等式构建审计框架,分离顺序敏感、不平衡与残余非局域性
- 实测发现17/18题对在直接评估下分布近确定性,标签映射显著影响结论
- 强调需先筛选测量饱和、反向平衡标签,再分析模型响应机制
人类调查数据中的问答顺序效应近似满足无参数的量子问卷(QQ)等式。本文将该等式发展为针对自回归大语言模型(LLM)序列二元判断的审计框架。理论上,刻画了满足QQ鲁棒性的机制家族,证明经典重复可精确复现该等式,并将QQ与秩-2上下文依赖准则结合,得到|q_QQ| ≤ OSS,从而区分顺序敏感性、QQ失衡与残余非局域性。方法上,提出受约束的多轮强制分支协议,通过反平衡标签映射和预设健康阈值,从下一词概率重构条件联合分布。初步实验在开放权重指令微调模型上显示核心测量问题:尽管所有健康阈值通过,但在直接评估框架下17/18题对、在角色框架下7/8题对的二元条件分布接近确定性。标签分配显著改变多个映射特定的QQ判断,且无任何题目被认证为残余上下文相关。因此,在测试条件下,观测到的QQ结果无法唯一识别响应机制,因存在饱和且标签敏感的测量接口。主要启示是方法论性的:若未建立充分分散性,不应将下一词概率解释为调查响应分布。故主张在分布级审计中,应优先进行饱和筛查与标签反向平衡。
原文摘要 · Abstract (English)
Question-order effects in human survey data have been reported to approximately satisfy the QQ (quantum question) equality, a parameter-free prediction of the standard projective quantum question-order model. We develop this equality into an audit framework for sequential binary judgments of autoregressive large language models (LLMs). Theoretically, we characterize mechanism families that satisfy QQ robustly, show that classical repetition can reproduce the equality exactly, and combine QQ with the rank-2 Contextuality-by-Default criterion through $|q_{QQ}| \le \mathrm{OSS}$. This separates order sensitivity, QQ imbalance, and residual contextuality rather than treating them as interchangeable signatures. Methodologically, we introduce a committed multi-turn forced-branch protocol that reconstructs order-conditioned joint distributions from next-token log-probabilities under counterbalanced label mappings and pre-specified health gates. A first-signal pilot on an open-weight instruction-tuned model reveals the central measurement problem. Although all pre-specified health gates passed, the binary-conditioned distributions were near-deterministic for 17 of 18 item pairs under the direct-evaluation framing and 7 of 8 under the persona framing. Label assignment materially changed several mapping-specific QQ verdicts, and no item was certified as residually contextual. Thus, under the tested conditions, the observed QQ outcomes did not uniquely identify a response mechanism in the presence of a saturated and label-sensitive measurement interface. The main implication is methodological: next-token probabilities should not be interpreted as survey-response distributions without first establishing adequate dispersion. We therefore argue that saturation screening and label counterbalancing should precede structural interpretation in distribution-level audits of LLM judgments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。