提出新评估框架XBOUND,从状态层面衡量设备控制智能体的真实能力。
XBOUND: Exploring Capability Boundaries of Device-Control Agents at the State Level
- 基于状态而非指令评估智能体行为,捕捉多交互场景下的真实表现
- 7B模型中UI-TARS最强,小模型在状态掌控上仍有明显短板
- 发现GPT规划是瓶颈,轨迹数据更利于统一指令理解
视觉语言模型的发展推动了设备控制智能体(DC agents)在图形用户界面(GUI)管理中的应用。随着系统复杂度提升,现有评估方法主要聚焦于指令级别,仅根据当前状态和历史执行记录判断动作,难以反映真实交互环境中的多样性。实际上,单个界面状态可能包含多个可操作控件,对应不同指令目标,形成多路径决策空间。为更全面评估智能体性能,本文提出状态级评估框架XBOUND,实现对每个状态中指令完成准确率的量化分析。实验揭示:在7B规模模型中,UI-TARS表现最优;当前智能体在指令统一性上呈现双峰分布;子7B模型仍受限于状态掌握能力。进一步分析表明,基于GPT的规划是关键瓶颈,而接地数据主要提升动作匹配效果,轨迹数据则更有利于指令统一。
原文摘要 · Abstract (English)
Recent advancements in vision-language models have increased interest in Device-Control Agents (DC agents) for managing graphical user interfaces (GUIs). With the growing complexity and integration of such agents into various applications, effective evaluation methods have become crucial. The current evaluation method for DC agents primarily focuses on the instruction level, providing the current state (e.g., screenshots) and past execution history to determine actions for target instructions, helping identify potential execution failures. However, in GUI environments, a single state may contain multiple interactive widgets, each linked to different instructions, presenting an opportunity for diverse actions based on various instruction targets. Evaluating the agent's performance solely at the instruction level may overlook the broader context of these interactions. To capture a more comprehensive view of agent performance, we propose a new evaluation method, XBOUND, to evaluate the accuracy of instruction completion on a per-state basis. XBOUND provides a state-level evaluation framework, serving as a tool to assess agents' capabilities within environmental states. Our evaluation yields several key insights: UI-TARS stands out as the strongest 7B model, current agents display a bimodal performance pattern in instruction unification, and sub-7B models remain limited in state mastery. We further identify GPT-based planning as a critical bottleneck, and show that grounding data mainly benefits action matching, while trajectory data is more effective for instruction unification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。