arXiv:2604.09574cs.AIcs.LG2026-04被引 3

让手机自动化工具学会模仿真人操作,躲过平台检测

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization

论文配图:Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization
图 1 · 摘自论文原文
  • 构建对抗性博弈框架,优化代理行为与人类的相似度
  • 新数据集显示普通AI操作易被识别,存在明显动作偏差
  • 提出可量化的人类化评估基准,适合安全合规研究者

自主移动GUI代理的兴起引发了数字平台的对抗性反制措施,但现有研究更关注任务完成能力与鲁棒性,忽视了抗检测这一关键维度。我们认为,为在以人为中心的生态系统中生存,代理必须具备人类化能力。本文提出“屏幕上的图灵测试”,将交互形式化为检测器与代理之间的最小最大优化问题,旨在最小化行为差异。我们收集了一个高保真移动触控动态数据集,并分析发现基于通用大模型的代理因运动学不自然而容易被识别。为此,我们建立了代理人类化基准(AHB)及检测指标,量化拟合度与任务性能间的权衡。通过启发式噪声到数据驱动的行为匹配方法,实验证明代理可在不牺牲性能的前提下实现理论与实证上的高拟合度。本工作将范式从‘能否完成任务’转向‘如何在人机共存环境中完成任务’,为对抗性数字环境中的无缝共存奠定基础。

原文摘要 · Abstract (English)

The rise of autonomous GUI agents has triggered adversarial countermeasures from digital platforms, yet existing research prioritizes utility and robustness over the critical dimension of anti-detection. We argue that for agents to survive in human-centric ecosystems, they must evolve Humanization capabilities. We introduce the ``Turing Test on Screen,'' formally modeling the interaction as a MinMax optimization problem between a detector and an agent aiming to minimize behavioral divergence. We then collect a new high-fidelity dataset of mobile touch dynamics, and conduct our analysis that vanilla LMM-based agents are easily detectable due to unnatural kinematics. Consequently, we establish the Agent Humanization Benchmark (AHB) and detection metrics to quantify the trade-off between imitability and utility. Finally, we propose methods ranging from heuristic noise to data-driven behavioral matching, demonstrating that agents can achieve high imitability theoretically and empirically without sacrificing performance. This work shifts the paradigm from whether an agent can perform a task to how it performs it within a human-centric ecosystem, laying the groundwork for seamless coexistence in adversarial digital environments.

人机交互自动化检测防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。