arXiv:2601.18305cs.CV2026-01中稿 · ACM MM 2026

让AI操作界面更像真人,通过模拟自然滑动手势提升交互成功率

SwipeGen: Bridging the Execution Gap in GUI Agents via Human-like Swipe Synthesis

  • 提出SwipeGen工具,生成多样且类人的滑动轨迹数据
  • 在复杂任务中使现有智能体滑动成功率最高提升2.46倍
  • 构建首个滑动交互评估基准SwipeBench,适合界面自动化研究者

尽管众多图形用户界面(GUI)智能体宣称可自动化用户交互任务,但至今仍鲜有能在真实场景中达到与人类相当的交互能力。通过实证分析,本文发现其根本原因在于滑动执行的僵化性:与人类对轨迹、速度和时机具有精细控制不同,现有智能体仅能执行简单、确定性的滑动行为,导致在复杂类人交互任务中频繁失败。由于缺乏开源的人类滑动训练数据,我们提出SwipeGen——首个合成多样化、类人滑动交互的工具,以及SwipeBench——首个评估智能体滑动交互质量的基准。大量实验表明,SwipeGen可使现有智能体的滑动执行成功率最高提升2.46倍。代码、数据集及模型已公开于https://github.com/TSKGHS17/SwipeGen。

原文摘要 · Abstract (English)

Despite numerous Graphical User Interface (GUI) agents claiming to automate user interaction tasks, to date, few achieve satisfactory interaction capability with human users in real-world scenarios. Through empirical analysis, this paper identifies the root cause of the limited interaction capability as the rigid swipe execution. In particular, unlike humans, who perform swipes with fine-grained control over trajectory, speed, and timing, existing agents can only conduct simplistic, deterministic swipe behaviors, leading to frequent failures on complicated user-like interaction tasks. Due to the lack of open-source human-like swipe training data, we propose SwipeGen, the first tool for synthesizing diverse and human-like swipe interactions, and SwipeBench, the first benchmark for evaluating agents' swipe interaction quality. Extensive experiments show that SwipeGen can improve the swipe execution success rate of existing agents by up to 2.46x. Our code, dataset, and model are available at https://github.com/TSKGHS17/SwipeGen.

GUI智能体滑动合成人机交互自动化测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。