首个面向短视频界面的动态屏幕智能体评测基准
Benchmarking Living-Screen-Native GUI Agents on Short-Video Platforms

- 构建动态屏幕环境,模拟短视频平台实时播放特性
- 模型普遍因观察过频或不足导致性能落后人类
- 适合研究交互智能体与观察策略的学者参考
当前GUI智能体假设屏幕状态在两次操作间保持不变,但短视频应用等真实界面内容持续播放,用户需自主决定观看时长。本文提出Living-Screen-Native GUI智能体新范式,设计LivingScreen基准,包含基于浏览器的真实环境、三层任务体系及兼顾准确率与信息效率的评估指标。对多个前沿模型评估发现,无一达到人类水平的成本-准确率表现,主要失败模式为过度或不足观察,揭示观察控制是未来智能体的关键缺失能力。所有数据与代码将公开于https://github.com/BITHLP/LivingScreen。
原文摘要 · Abstract (English)
GUI agents today assume a static screen, where the world is frozen between two actions. However, real interfaces such as short-video applications violate this assumption, as their content keeps playing, and a competent user must decide what to watch and for how long. We formalize this task as Living-Screen-Native GUI agents and introduce LivingScreen, the first benchmark instantiating it on short-video platforms, with a faithful browser-based environment, a three-tier task suite, and metrics that jointly score accuracy and information efficiency. Evaluating extensive frontier models, we find that none reaches the human cost-accuracy performance, and that their dominant failure mode is over- and under-observation, pointing to observation control as a missing capability axis for future GUI agents. All data and code will be available at https://github.com/BITHLP/LivingScreen.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。