解决图形界面操作中语义与执行的精度鸿沟问题。
PAGER: Bridging the Semantic-Execution Gap in Point-Precise Geometric GUI Control

- 构建依赖结构化规划与像素级执行的拓扑感知代理
- 在4906个任务上实现62%步骤成功率,任务成功率达4.1倍提升
- 适合需要精准界面操作的研究者与工业应用开发者
大型视觉语言模型虽推动了GUI智能体发展,但主要依赖宽松的区域容忍范式。当任务要求点级精度时,局部坐标误差会引发连锁拓扑错误,导致最终结果失效。我们提出精度敏感型GUI任务新范式,需点级准确、几何感知验证及抗依赖误差传播能力。为此引入PAGE Bench基准,包含4,906个问题和超过224,000条像素级操作标注。提出PAGER,通过依赖结构化规划与像素级执行实现拓扑感知控制;采用像素接地监督微调建立可执行动作语法,结合状态条件化的几何反馈进行精度对齐强化学习,缓解训练-推理偏差。实验显示:通用多模态模型动作类型准确率超88%,但任务成功率低于6%;而PAGER使任务成功率提升4.1倍,步骤成功率从不足9%跃升至62%,确立点精确GUI控制新基准。
原文摘要 · Abstract (English)
Large vision-language models have significantly advanced GUI agents, enabling executable interaction across web, mobile, and desktop interfaces. Yet these gains largely rely on a forgiving region-tolerant paradigm, where many nearby pixels inside the same component remain valid. Precise geometric construction breaks this assumption: actions must land on points in continuous canvas space rather than tolerant regions. Because geometric primitives carry ontological dependencies, a local coordinate error can induce cascading topological failures that distort downstream objects and invalidate the final construction. We identify this regime as precision-sensitive GUI tasks, requiring point-level accuracy, geometry-aware verification, and robustness to dependency-driven error propagation. To benchmark it, we introduce PAGE Bench, with 4,906 problems and over 224K process-supervised, pixel-level GUI actions. We further propose PAGER, a topology-aware agent that decomposes construction into dependency-structured planning and pixel-level execution. Pixel-grounded supervised tuning establishes executable action grammar, while precision-aligned reinforcement learning mitigates rollout-induced exposure bias through state-conditioned geometric feedback. Experiments reveal a pronounced Semantic-Execution Gap: general multimodal models can exceed 88% action type accuracy yet remain below 6% task success. PAGER closes this gap, delivering 4.1x higher task success than the strongest evaluated general baseline and raising step success rate from below 9% for GUI-specialized agents to over 62%, establishing a new state of the art for point-precise GUI control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。