arXiv:2606.06322cs.AI2026-06

构建首个拖拽交互基准数据集,推动GUI自动化任务发展

DragOn: A Benchmark and Dataset for Drag-Based GUI Interactions

论文配图:DragOn: A Benchmark and Dataset for Drag-Based GUI Interactions
图 1 · 摘自论文原文
  • 构建跨四领域的拖拽交互数据集,涵盖文本高亮等场景
  • 包含286万张训练图与350万任务,评估集2000例
  • 验证主流模型在复杂拖拽任务上的提升潜力

GUI智能体——基于视觉的模型,通过图形用户界面控制桌面、网页浏览器和移动设备——有望自动化大量数字任务。尽管已有百万级数据集推动点击定位取得进展,但拖拽定位(如拖拽、滑动、高亮)的数据仍小一个数量级,现有模型在复杂拖拽操作上表现不足。我们提出DragOn,一个覆盖四个领域(文本高亮、单元格选择、元素缩放、滑块操作)的拖拽接地基准与训练数据集。该数据集包含286万张训练截图和350万条训练任务,以及一个2000样本的独立评估集。我们评估了商用模型(GPT、Claude)和开源模型(Qwen、Kimi、Holo),以及在我们的训练数据上微调的Qwen VLM。结果表明,该数据集可显著提升当前先进模型在下游计算机使用任务中的表现。

原文摘要 · Abstract (English)

GUI agents - vision-based models that control desktops, web browsers, and mobile devices through graphical user interfaces - promise to automate a wide range of digital tasks. While million-scale datasets have enabled substantial progress on click-grounding, drag grounding (e.g. drag-and-drop, swipe, highlight) data remains an order of magnitude smaller and current models fall short on complex drag-based interactions. We introduce DragOn, a drag grounding benchmark and training dataset covering four domains: text highlighting, cell selection, element resizing and slider manipulation. The dataset comprises 286K training screenshots and 3.5M training tasks, plus a 2000-example held-out evaluation suite. We evaluate proprietary (GPT, Claude) and open-weight (Qwen, Kimi, Holo) models, as well as a Qwen VLM fine-tuned on our training data. Results suggest that our dataset could improve performance of state-of-the-art models on downstream computer-use tasks.

GUI智能体拖拽识别基准测试视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。