arXiv:2410.11871cs.HCcs.AI2024-10中稿 · Interspeech 2025被引 15

用小模型实现精准界面操作,低成本高效自动化。

TinyClick: Single-Turn Agent for Empowering GUI Automation

  • 基于视觉语言模型,直接定位用户指令对应的界面元素坐标。
  • 在Screenspot和OmniAct上表现优异,仅0.27B参数,延迟极低。
  • 训练只需56 GPU小时,适合资源有限的研究者快速实验。

我们提出一个用于用户界面交互任务的UI智能体,基于视觉语言模型Florence-2-Base。该智能体核心任务是根据用户指令准确识别屏幕中对应界面元素的坐标。在Screenspot和OmniAct标注数据集上表现出色,模型参数量仅为0.27B,延迟极低。训练仅需56 GPU小时(约40美元),显著降低计算成本。性能提升得益于视觉专用多任务训练与基于多模态大模型的数据增强策略。我们希望减少对昂贵算力和人工标注数据的依赖,推动更包容、可持续的UI智能体研究。

原文摘要 · Abstract (English)

We present an UI agent for user interface (UI) interaction tasks, using Vision-Language Model Florence-2-Base. The agent's primary task is identifying the screen coordinates of the UI element corresponding to the user's command. It demonstrates very strong performance on Screenspot and OmniAct annotations, while maintaining a very small size of 0.27B parameters and minimal latency. Moreover, training needs small compute budget of 56 GPU-hours (worth about 40 USD). Relevant improvement comes from vision-specific multi-task training and MLLM-based data augmentation. We hope that decreased needs for expensive compute resources and manually annotated data will allow to facilitate more inclusive and sustainable research of UI agents.

GUI自动化小模型视觉语言模型低成本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。