提出协议级可识别性审计,验证大模型推理评估是否真正测量目标行为。
Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation

- 构建协议级可识别性审计框架,检验观测能否区分不同策略。
- 基线观测将7个策略合并为1类,全支持可分离出7类且无误判。
- 适用于评估大模型干预响应真实性的研究者,尤其关注评估设计有效性。
LLM基准分数可能在观测协议无法识别其意图测量的行为属性时仍保持精确。在受控的求解器基础设置中,我们针对有限行为策略类定义了协议级可识别性审计:给定策略集H、观测支持O和估计量τ,测试O是否能区分所有τ值不同的策略对。该审计无需模型调用,解决了诊断案例:仅基础观测将7个冻结确定性策略合并为一个等价类;完整支持可区分7个类别且无跨估计量冲突;每个剔除一项的支持均保留构造性冲突证据。实证显示,两种受限生成变体的成对有效性均为1.0,但基础准确率与选择性响应保真度差异显著——六组平衡奥数转换方向上分别为0.620与0.324(置信区间[0.600, 0.642]与[0.304, 0.345]),第二组确定性源中差距再现(0.646 vs. 0.331)。审计还合成最小识别支持O*:仅需2个单元而非完整的36单元张量。此案例表明,评估设计的有效性可在模型推理前结构化验证,且基础正确性不等于干预响应保真度。
原文摘要 · Abstract (English)
LLM benchmark scores can be precise even when the observation protocol does not identify the behavioral property they are intended to measure. In a controlled, solver-grounded setting, we formalize a protocol-level identifiability audit over a finite behavioral policy class: given policies H, observation support O, and estimand $τ$, we test whether O separates every pair with different $τ$. The audit requires zero model calls and resolves our diagnostic case: base-only observation collapses seven frozen deterministic policies into one equivalence class; full support yields seven classes and no cross-estimand collisions; every leave-one-out support retains a constructive collision witness. Empirically, both constrained-generation variants have pair-validity 1.0, yet base accuracy and selective-response fidelity diverge - 0.620 versus 0.324 across six balanced oracle-transition directions (cluster-bootstrap 95% CI [0.600, 0.642] vs. [0.304, 0.345]) - and the gap recurs on a second deterministic source (0.646 vs. 0.331). The audit also synthesizes a minimum identifying support $O^*$ for the frozen policy class: two cells instead of the full 36-cell tensor. This case shows how evaluation-design validity can be checked structurally before model inference and why base correctness does not determine intervention-response fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。