评估大模型决策步骤贡献度,发现现有方法不可靠。
Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay
- 用执行回放构建条件真值,检验各信用信号有效性
- 多数信号与真实贡献相关性极低,部分结果不显著
- 适合关注模型可解释性与评估方法论的研究者
在单智能体工具环境 ALFWorld 中,基于执行回放的策略条件真值,我们审计了多种步骤级信用信号:LLM 判分、结果条件对数似然比、策略自身置信度。结果显示,这些信号在匹配边缘分布的随机对照下均无可靠增量保真度。校正回放目标可靠性后,隐式信用保真度接近零,判分信用则无法得出结论。现有评估以步骤正确性为标准,而我们以步骤贡献(重采样策略替代路径对结果的影响)为基准,二者明显分离。真值结构显示,在定义的决策点中30.5%具有非零回放对比,且可测性依赖模型——两个相似规模策略的无支持反事实点比例相差一倍(13.1% vs. 26.8%)。失败模式清晰:隐式信用反映策略流畅度(中位秩相关+0.75),结果条件未提供因果信息(偏相关-0.004,Qwen)。仅依赖置信度的路由在关键步骤上表现随机,但每回合节省13.1%判分成本(每轨迹14.0%)。在七臂预注册训练实验中,无一训练组显著优于未训练策略,检查点的仪器特征与有效训练剂量一致——优化步数跨度达数量级差异,而非信用内容。信用规则比较必须匹配有效样本量,否则测量的是剂量而非信用。
原文摘要 · Abstract (English)
Audited against policy-conditional ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level credit signals we audit -- LLM-judge scores, outcome-conditioned logprob ratios, or the policy's own confidence -- shows reliable incremental fidelity beyond its own marginal-matched shuffled control. Correcting for replay-target reliability leaves implicit fidelity bounded near zero and judge fidelity inconclusive at the achieved target reliability. Existing evaluations grade these signals against annotated step *correctness*; we audit them against step *contribution* -- what re-sampling the policy's own alternatives at each decision point and rolling forward changes about the outcome -- and they come apart. The ground truth is structured: 30.5% of decision points where it is defined exhibit a nonzero replay contrast at the achieved sampling resolution, and measurability is model-dependent -- the fraction of points with no policy-supported counterfactual differs twofold (13.1% vs. 26.8%) between two similar-scale policies. The failure mode is identifiable: implicit credit echoes the policy's fluency (median rank correlation +0.75, replicating at +0.70 in a second family under a corrected instrument), while outcome conditioning adds no causal information (partial correlation -0.004, Qwen). A confidence-only router recovers pivotal steps at chance level, but cuts judge cost by 13.1% per turn (14.0% per trajectory). In a seven-arm pre-registered training experiment, no arm reliably outperforms the untrained policy, and the checkpoints' apparent instrument signature is statistically consistent with mediation by effective training dose in this design -- sparser credit retains fewer examples, an order-of-magnitude spread in optimizer steps -- not credit content. Comparisons of credit rules must match effective sample size, or they measure dose, not credit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。