arXiv:2605.01847cs.AI2026-05

评测大模型任务承诺一致性,发现成功不等于可靠。

NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles

论文配图:NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
图 1 · 摘自论文原文
  • 用人工校准的探针测试模型是否坚持任务承诺。
  • 32个模型中31个按承诺度排名变化,且更抗干扰。
  • 适合关注模型可信性与长期一致性的研究者。

现有仅看结果的评估方式无法判断智能体在多轮任务中是否保持必要承诺。NeuroState-Bench 是一个通过基准定义的探针问题进行人工校准的评测集,包含144个确定性任务和306个侧向探针,覆盖八类认知失效模式,含干净与干扰变体、三个难度等级。主评估涵盖32个模型,其中16个本地模型与16个大型模型托管版本,使用相同评测流程。人工校准基于104个采样任务单元、216条原始标注和108条仲裁任务行,加权卡帕系数0.977,组内相关系数ICC(2,1)=0.977。实证显示任务成功率与承诺完整性存在明显分离:成功率最高者并非承诺度最高者,31/32模型在以承诺度取代成功率后排名改变,且承诺度排序对干扰更稳定。核心无置信分数HCCIS-CORE在诊断终局任务失败上达到0.8469 AUC和0.6992 PR-AUC;传统全启发式版本HCCIS-FULL为0.7997 AUC和0.6410 PR-AUC。探针准确率与状态漂移指标在ROC-AUC(0.8587)和布里尔/期望误差上表现更优,但HCCIS-CORE在点估计PR-AUC上更高,且与评测目标构念关联更紧密。神经增强变体HCCIS+N整体表现较弱,随机子空间控制接近随机水平。因此,NeuroState-Bench 提供了比原局部子集更广范围的承诺失败暴露能力。

原文摘要 · Abstract (English)

Outcome-only evaluation under-specifies whether an evaluated agent profile preserves the commitments required to solve a multi-turn task coherently. NeuroState-Bench is a human-calibrated benchmark that operationalizes commitment integrity through benchmark-defined side-query probes rather than inferred hidden activations. The released inventory contains 144 deterministic tasks and 306 benchmark-defined side-query probes spanning eight cognitively motivated failure families, paired clean and distractor variants, and three difficulty bands. The main 32-profile evaluation contains a fixed 16-profile local subset and a matched 16-profile hosted large-model subset evaluated through the same benchmark pipeline. Human calibration uses the final merged reporting scope: 104 sampled task units, 216 raw annotations, and 108 adjudicated task rows, with weighted kappa = 0.977 and ICC(2,1) = 0.977. Empirically, task success and commitment integrity diverge across this expanded grid: the success leader is not the integrity leader, 31 of 32 profiles change rank when integrity replaces task success, and integrity rankings are more stable under distractor perturbation. The primary confidence-free score HCCIS-CORE reaches 0.8469 AUC and 0.6992 PR-AUC for post-probe diagnostic discrimination of terminal task failure; the legacy full heuristic variant HCCIS-FULL reaches 0.7997 AUC and 0.6410 PR-AUC. Probe accuracy and state drift achieve slightly higher ROC-AUC, 0.8587, and better Brier/ECE, while HCCIS-CORE has substantially higher point-estimate PR-AUC and remains more closely tied to the benchmark's intended construct. The exploratory neural-augmented variant HCCIS+N is weaker overall, and a randomized subspace control approaches chance. NeuroState-Bench therefore contributes a calibrated evaluation axis for exposing commitment failures over a broader model grid than the original local-only subset.

大模型评估承诺一致性基准测试可信智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。