用状态化界面结构提升界面智能体的理解与预测能力
ScreenLLM: Stateful Screen Schema for Efficient Action Understanding and Prediction
- 构建动态界面状态图谱,捕捉用户操作意图与行为序列
- 在开源与私有数据集上实现高精度动作预测与行为建模
- 适合开发智能自动化工具和个性化用户助手的开发者
图形用户界面(GUI)智能体是能够解析与生成操作的自主系统,可实现智能用户辅助与自动化。其有效训练面临监督信号稀疏、大规模数据扩展性差以及深层用户理解需求等挑战。本文提出状态化屏幕结构,一种高效表示GUI交互的方式,能随时间捕捉关键用户行为与意图。基于此,我们构建ScreenLLM,一组针对高级界面理解与动作预测优化的多模态大语言模型。在开源及专有模型上的大量实验表明,ScreenLLM能准确建模用户行为并预测下一步操作。本工作为可扩展、鲁棒且智能的GUI智能体奠定基础,显著提升多样软件环境中的用户交互体验。
原文摘要 · Abstract (English)
Graphical User Interface (GUI) agents are autonomous systems that interpret and generate actions, enabling intelligent user assistance and automation. Effective training of these agent presents unique challenges, such as sparsity in supervision signals, scalability for large datasets, and the need for nuanced user understanding. We propose stateful screen schema, an efficient representation of GUI interactions that captures key user actions and intentions over time. Building on this foundation, we introduce ScreenLLM, a set of multimodal large language models (MLLMs) tailored for advanced UI understanding and action prediction. Extensive experiments on both open-source and proprietary models show that ScreenLLM accurately models user behavior and predicts actions. Our work lays the foundation for scalable, robust, and intelligent GUI agents that enhance user interaction in diverse software environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。