arXiv:2606.04701cs.CVcs.CL2026-06

首个面向短视频界面的动态屏幕智能体评测基准

Benchmarking Living-Screen-Native GUI Agents on Short-Video Platforms

论文配图:Benchmarking Living-Screen-Native GUI Agents on Short-Video Platforms
图 1 · 摘自论文原文
  • 构建动态屏幕环境,模拟短视频平台实时播放特性
  • 模型普遍因观察过频或不足导致性能落后人类
  • 适合研究交互智能体与观察策略的学者参考

当前GUI智能体假设屏幕状态在两次操作间保持不变,但短视频应用等真实界面内容持续播放,用户需自主决定观看时长。本文提出Living-Screen-Native GUI智能体新范式,设计LivingScreen基准,包含基于浏览器的真实环境、三层任务体系及兼顾准确率与信息效率的评估指标。对多个前沿模型评估发现,无一达到人类水平的成本-准确率表现,主要失败模式为过度或不足观察,揭示观察控制是未来智能体的关键缺失能力。所有数据与代码将公开于https://github.com/BITHLP/LivingScreen。

原文摘要 · Abstract (English)

GUI agents today assume a static screen, where the world is frozen between two actions. However, real interfaces such as short-video applications violate this assumption, as their content keeps playing, and a competent user must decide what to watch and for how long. We formalize this task as Living-Screen-Native GUI agents and introduce LivingScreen, the first benchmark instantiating it on short-video platforms, with a faithful browser-based environment, a three-tier task suite, and metrics that jointly score accuracy and information efficiency. Evaluating extensive frontier models, we find that none reaches the human cost-accuracy performance, and that their dominant failure mode is over- and under-observation, pointing to observation control as a missing capability axis for future GUI agents. All data and code will be available at https://github.com/BITHLP/LivingScreen.

GUI智能体动态屏幕评测基准短视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。