arXiv:2606.17541cs.LGcs.AI2026-06被引 1

用轨迹偏好评估替代成功判断,显著减少评测中的平局现象。

Offline Preference-Based Trajectory Evaluation

  • 直接比较轨迹的进展与完成时间,而非仅看最终成败
  • 平局率从75%降至35%,提升评测区分度和效率
  • 适合评估复杂智能体系统,尤其在数据稀缺时更有效

离线评估智能体系统时常将轨迹简化为最终成败,忽略中间进展,导致大量平局,严重降低样本有效性和区分能力。本文提出基于偏好的轨迹评估方法,通过对比轨迹在进展速度与返奖时间上的偏好关系进行评价。在多个智能体与交互基准测试中,传统成功指标约75%的实例出现平局,而轨迹感知的偏好评估将平局率降至约35%,显著提升判别力、排序稳定性和数据利用效率。结果表明,评测饱和现象可能不仅源于数据收集不足或任务难度,也与评估方式选择有关。

原文摘要 · Abstract (English)

Offline evaluation of agentic systems often collapses trajectories to terminal success, discarding information about partial progress and inducing widespread ties, creating substantial statistical inefficiency by reducing effective sample size and weakening the ability to distinguish systems. We propose preference-based trajectory evaluation, which compares trajectories directly through temporal preferences over progress and time-to-return profiles. We find that, across diverse agentic and interactive benchmarks, standard success-based metrics produce tied comparisons on roughly 75% of instances, whereas trajectory-aware preferences reduce ties to roughly 35%, improving discriminative power, ranking stability, and data efficiency. Our results suggest that benchmark saturation, often attributed to poor data collection or problem difficulty, may also be explained by the choice of evaluation measure.

智能体评估轨迹评价偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。