测试UI智能体在不同界面变化下的可靠性,发现表现波动超50%。
OpenApps: Simulating Environment Variations to Measure UI-Agent Reliability
- 构建可配置的六款应用生态,生成数千种界面变体。
- 10,000次评估显示,多模态智能体成功率在不同界面下波动超50%。
- 适合评估部署前智能体鲁棒性,尤其关注界面适应能力。
可靠性是实现自主UI智能体潜力的关键,这类多模态智能体以人类方式直接与应用交互,用户必须信任其完成任务的能力。当前评估依赖固定环境,通常是现有应用的克隆,仅能反映智能体在特定环境中的表现。但实际部署时,应用设计与内容的变化可能影响任务完成能力。为弥补这一盲区,我们开发了OpenApps——一个轻量级开源生态系统,包含六款可配置外观与内容的应用(如消息、日历、地图等),仅需单个CPU即可运行,支持快速生成并部署数千个应用版本。我们对七种领先多模态智能体进行了超过10,000次独立评估,发现虽然固定应用内的可靠性相对稳定,但在跨应用变化下表现差异巨大:许多智能体的任务成功率在不同变体间波动超过50%。例如,Kimi-VL-3B在所有任务上的平均成功率在不同版本间从63%降至4%。我们还观察到智能体行为(如循环执行或虚构操作)随环境配置显著变化。这些发现凸显了在应用变化维度上测量可靠性的必要性。OpenApps已开源:https://facebookresearch.github.io/OpenApps/
原文摘要 · Abstract (English)
Reliability is key to realizing the promise of autonomous UI-Agents, multimodal agents that directly interact with apps in the same manner as humans, as users must be able to trust an agent to complete a given task. Current evaluations rely on fixed environments, often clones of existing apps, which are limited in that they can only shed light on whether or how often an agent can complete a task within a specific environment. When deployed however, agents are likely to encounter variations in app design and content that can affect an agent's ability to complete a task. To address this blind spot of measuring agent reliability across app variations, we develop OpenApps, a light-weight open-source ecosystem with six apps (messenger, calendar, maps, etc.) that are configurable in appearance and content. OpenApps requires just a single CPU to run, enabling easy generation and deployment of thousands of versions of each app. Specifically, we run more than 10,000 independent evaluations to study reliability across seven leading multimodal agents. We find that while standard reliability within a fixed app is relatively stable, reliability can vary drastically when measured across app variations. Task success rates for many agents can fluctuate by more than $50\%$ across app variations. For example, Kimi-VL-3B's average success across all tasks fluctuates from $63\%$ to just $4\%$ across app versions. We also find agent behaviors such as looping or hallucinating actions can differ drastically depending on the environment configuration. These initial findings highlight the importance of measuring reliability along this new dimension of app variations. OpenApps is available at https://facebookresearch.github.io/OpenApps/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。