arXiv:2511.20067cs.AIcs.HC2025-11中稿 · AAAI被引 3

用视觉语言模型自动判断电脑操作任务是否完成,提升智能代理的自纠错能力。

"Are We Done Yet?": A Vision-Based Judge for Autonomous Task Completion of Computer Use Agents

  • 通过截图和任务描述,用视觉语言模型判断任务是否完成。
  • 任务成功检测准确率达73%,反馈使整体成功率提升27%。
  • 适合开发自动化桌面助手或需要自我评估的AI代理的人参考。

计算机使用代理(CUAs)旨在自主操作数字界面,但常无法可靠判断任务是否完成。本文提出一种基于视觉的自主评估与反馈框架,利用视觉语言模型直接从截图和任务描述中判断任务完成情况。构建了涵盖42个macOS内置应用、1,260个经人工标注任务的大型数据集。实验显示,该框架在任务成功检测上最高达73%准确率;当引入评估反馈后,整体任务成功率平均提升27%。结果表明,视觉评估可作为有效反馈机制,显著增强自主计算机使用代理的可靠性与自修正能力。

原文摘要 · Abstract (English)

Computer Use Agents (CUAs) are designed to autonomously operate digital interfaces, yet they often fail to reliably determine whether a given task has been completed. We present an autonomous evaluation and feedback framework that uses vision-language models to assess task completion directly from screenshots and task descriptions. Our dataset covers 42 built-in macOS applications and 1,260 human-labeled tasks across a wide range of scenarios. Our framework achieves up to 73 percent accuracy in task success detection and yields an average relative improvement of 27 percent in overall task success when evaluator feedback is applied. These results show that vision-based evaluation can serve as an effective feedback mechanism that improves the reliability and self-correction of autonomous computer-use agents.

任务评估视觉语言模型自动化代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。