用视频评估计算机操作智能体,准确率超84%
Video-Based Reward Modeling for Computer-Use Agents
- 通过关键帧视频建模任务完成度,不依赖内部推理过程
- 在53000条视频数据上训练,准确率达84.7%,召回87.7%
- 适合需要高效、通用评估智能体行为的研究者
计算机使用智能体(CUAs)能力日益增强,但难以规模化评估其轨迹是否真正满足用户指令。本文研究从执行视频中进行奖励建模:一种独立于智能体内部推理与动作的键帧序列。尽管该方法与具体模型无关,仍面临布局高度冗余、成功线索微弱且局部等挑战。我们构建了包含53,000个高质量视频-任务-奖励三元组的ExeVR-53k数据集,并提出对抗性指令转换以生成带步骤级标注的负样本。为支持长时、高分辨率视频学习,设计时空标记剪枝策略,移除同质区域和持续存在标记,保留决定性界面变化。基于上述组件,微调得到仅需用户指令与视频序列即可预测任务成功的执行视频奖励模型(ExeVRM)。ExeVRM 8B在Ubuntu、macOS、Windows和Android上取得84.7%准确率与87.7%召回率,优于GPT-5.2和Gemini-3 Pro等强基线模型,且具备更精确的时间定位能力。结果表明,视频执行奖励建模可作为可扩展、模型无关的CUA评估工具。
原文摘要 · Abstract (English)
Computer-using agents (CUAs) are becoming increasingly capable; however, it remains difficult to scale evaluation of whether a trajectory truly fulfills a user instruction. In this work, we study reward modeling from execution video: a sequence of keyframes from an agent trajectory that is independent of the agent's internal reasoning or actions. Although video-execution modeling is method-agnostic, it presents key challenges, including highly redundant layouts and subtle, localized cues that determine success. We introduce Execution Video Reward 53k (ExeVR-53k), a dataset of 53k high-quality video--task--reward triplets. We further propose adversarial instruction translation to synthesize negative samples with step-level annotations. To enable learning from long, high-resolution execution videos, we design spatiotemporal token pruning, which removes homogeneous regions and persistent tokens while preserving decisive UI changes. Building on these components, we fine-tune an Execution Video Reward Model (ExeVRM) that takes only a user instruction and a video-execution sequence to predict task success. Our ExeVRM 8B achieves 84.7% accuracy and 87.7% recall on video-execution assessment, outperforming strong proprietary models such as GPT-5.2 and Gemini-3 Pro across Ubuntu, macOS, Windows, and Android, while providing more precise temporal attribution. These results show that video-execution reward modeling can serve as a scalable, model-agnostic evaluator for CUAs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。