首个评估计算机使用智能体奖励模型的基准,解决传统方法评估难问题。
CUARewardBench: A Benchmark for Evaluating Reward Models on Computer-using Agent
- 构建首个涵盖结果与过程奖励模型的评测基准,支持全流程评估。
- 覆盖10类软件、7种智能体架构,成功率达25.9%-50.8%,数据经专家严格标注。
- 提出统一投票集成方法,精度超单模型30%以上,适合训练与评估奖励模型者使用。
计算机使用智能体(CUAs)通过自然交互操作系统与软件界面完成任务。尽管脚本验证器被广泛采用,但其可扩展性差且无法进行步骤级评估。奖励模型提供了有前景的替代方案,但其在CUA评估中的有效性仍缺乏系统研究。为此,我们提出CUARewardBench,包含四大贡献:(1) 首个全面的CUA奖励模型评测基准,首次对结果奖励模型(ORM)和过程奖励模型(PRM)进行轨迹级与步骤级评估;(2) 多样化、实用且可靠的高质量数据集,涵盖10类软件与7种智能体架构,成功率25.9%-50.8%,所有轨迹经精心设计协议与严格质量控制标注;(3) 全面分析揭示当前奖励模型关键缺陷,包括视觉推理能力不足、知识缺失,以及通用视觉语言模型(VLM)优于专用模型;(4) 提出一致提示集成(UPE),通过严格一致投票与策略性提示配置,显著提升可靠性,实现ORM 89.8%精确率与93.3%负预测值,PRM达81.7%精确率与85.1%负预测值,显著优于单模型与传统集成方法。
原文摘要 · Abstract (English)
Computer-using agents (CUAs) enable task completion through natural interaction with operating systems and software interfaces. While script-based verifiers are widely adopted for evaluation, they suffer from limited scalability and inability to provide step-wise assessment. Reward models offer promising alternatives, but their effectiveness on CUA evaluation remains largely underexplored. To address this gap, we present CUARewardBench, comprising four key contributions: (1) First-ever Comprehensive CUA Reward Benchmark: We introduce the first benchmark for evaluating both outcome reward models (ORM) and process reward models (PRM) on CUA tasks, enabling systematic assessment across trajectory-level and step-level evaluation. (2) Diverse, Practical and Reliable Dataset: CUARewardBench encompasses trajectories from 10 software categories and 7 agent architectures with varying performance levels (25.9%-50.8% success rates). All trajectories are expertly annotated through carefully designed protocols, with rigorous quality control to ensure reliability and practical applicability. (3) Comprehensive Analysis and Insights: Through extensive experiments across 7 vision-language models and 3 prompt templates, we reveal critical limitations of current CUA RMs, including insufficient visual reasoning capabilities, knowledge deficiencies, and the superiority of general VLMs over specialized CUA models for reward evaluation. (4) Unanimous Prompt Ensemble (UPE): Based on the insights from our comprehensive analysis, we propose UPE, a novel ensemble method that significantly enhances reward model reliability through strict unanimous voting and strategic prompt-template configurations. UPE achieves 89.8% precision and 93.3% NPV for ORM, and 81.7% precision and 85.1% NPV for PRM, substantially outperforming single VLMs and traditional ensemble approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。