arXiv:2608.17713cs.LG2026-08

交叉视图对应关系影响智能体评估结果,需验证并传播其不确定性。

Cross-View Correspondence Is a Measurement Intervention: Two-Sided Validation for Agent Evaluation and Credit Assignment

  • 提出双向验证框架,确保去噪与响应保留的合理性
  • 55.9%的轨迹对在不同回溯中时间定位不一致,显示信用分配不可靠
  • 适用于需要可信评估与责任归属的自动化系统研发

智能体评估与基于轨迹的学习常通过后处理对应关系比较不同视角的输出,但该对应关系实为测量干预:省略它会制造敏感性,过度优化的映射会制造虚假不变性,多个最优对应会导致机制标签和学习信用无法确定。本文构建有效性理论与审计体系,包含三部分:干扰消除与响应保留的双向验证、下游结论的所有最优解识别、有效性确立后的不确定性传播。我们刻画了保持响应的干扰消除线性可行性边界,计算精确最优对应集的紧致范围,并给出无需分布假设的证书:仅当所有最优解在某信用坐标上符号一致时才保留其非零值。在公开代码与SQL管道中,两个确定性最优回溯对1,586个非零轨迹对中55.9%的时间定位不一致;两个冻结的800次回放审计(包括任务与种子不重叠的复现)暴露了意图的逐轮信用反转,尽管一个干净的快速启动子集未见此现象。预先注册的传输门在自然响应上失败;冻结修正与留出控制组表明,仅在良性样本上校准的映射会抹除所有保留的有害响应,而双向验证则选出保持响应的替代方案。因此,交叉视图对应必须声明、验证并传播不确定性,才能支撑可靠的评估与信用分配。

原文摘要 · Abstract (English)

Agent evaluations and trace-based learning often compare outputs across transformed views through a post-response correspondence treated as neutral preprocessing. We show that this correspondence is a measurement intervention: omitting it can manufacture sensitivity, an over-aggressive map can manufacture invariance, and multiple optimal correspondences can leave mechanism labels and signed learning credit unidentified. We develop a validity theory and audit with three components: two-sided validation of nuisance removal and response preservation, all-optima identification of downstream conclusions, and uncertainty propagation after validity is established. We characterize the linear feasibility boundary for response-preserving nuisance removal, compute sharp ranges over exact-optimum correspondence sets, and give a distribution-free certificate that retains a credit coordinate only when all exact optima agree on its nonzero sign. Across public code and SQL pipelines, two deterministic optimal tracebacks disagree on temporal localization for 55.9% of 1,586 nonzero trajectory pairs; two frozen 800-rollout tool-use audits, including a task-and-seed-disjoint replication, expose exact-optimum reversals of intended turn-level credit, although a clean public quick-start subset shows none. A pre-registered transport gate failed on natural responses; frozen corrected and held-out controls then show that a map calibrated only on benign examples erases every retained harmful response, while two-sided validation selects response-preserving alternatives. Cross-view correspondence must therefore be declared, validated, and propagated into uncertainty before agent evaluation or credit assignment supports a point conclusion.

智能体评估信用分配不确定性传播

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。