提出文本拖拽数据集与评测基准,推动GUI交互从点击扩展到拖拽。
Beyond Clicking:A Step Towards Generalist GUI Grounding via Text Dragging
- 构建16.1万条文本拖拽数据,支持大规模训练
- 在5333个样本上实现拖拽性能显著提升
- 适用于开发更通用的GUI智能体
图形用户界面(GUI)定位是将人类指令映射为GUI操作的基础,现有模型在点击类任务上表现良好,但对更常见的拖拽操作研究不足。为此,我们提出GUI-Drag,一个通过可扩展流水线合成的16.1万条文本拖拽数据集;并构建ScreenDrag基准,包含5333个样本,覆盖三类界面上下文,配套三项专用于评估文本拖拽能力的指标。基于GUI-Drag采用高效持续训练策略的模型,在ScreenDrag上实现显著性能提升,同时保持在ScreenSpot、ScreenSpot-v2和OSWorld-G上的原有点击性能。本工作推动了超越点击的通用化GUI定位研究,为构建真正通用的GUI智能体铺路。所有数据、代码与模型均已开源。
原文摘要 · Abstract (English)
Graphical user interface (GUI) grounding, the process of mapping human instructions to GUI actions, serves as a fundamental basis to autonomous GUI agents. While existing grounding models achieve promising performance to simulate the mouse click action on various click-based benchmarks, another essential mode of mouse interaction, namely dragging, remains largely underexplored. Yet, dragging the mouse to select and manipulate textual content represents a prevalent and important usage in practical GUI scenarios. To narrow this gap, we first introduce GUI-Drag, a diverse dataset of 161K text dragging examples synthesized through a scalable pipeline. To support systematic and robust evaluation, we further construct ScreenDrag, a benchmark with 5,333 examples spanning three levels of interface context, together with three dedicated metrics designed for assessing text dragging capability. Models trained on GUI-Drag with an efficient continual training strategy achieve substantial improvements on ScreenDrag, while preserving the original click-based performance on ScreenSpot, ScreenSpot-v2, and OSWorld-G. Our work encourages further research on broader GUI grounding beyond just clicking and paves way toward a truly generalist GUI grounding model. All benchmark, data, checkpoints, and code are open-sourced and available at https://osu-nlp-group.github.io/GUI-Drag.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。