arXiv:2601.19258cs.HCcs.AI2026-01中稿 · ACM CHI Conference…被引 2

构建数据集助力识别手机界面隐藏操作,提升自动化效率。

GhostUI: Unveiling Hidden Interactions in Mobile UI

论文配图:GhostUI: Unveiling Hidden Interactions in Mobile UI
图 1 · 摘自论文原文
  • 构建包含前后截图与手势元数据的GhostUI数据集
  • 在预测隐藏操作和交互后界面时准确率显著提升
  • 适合研究移动端自动化与智能助手的开发者使用

现代移动应用依赖无视觉提示的隐藏交互(如长按、滑动)来实现功能,避免界面杂乱。虽然资深用户可通过使用经验或引导教程发现这些操作,但其隐性特征使大多数用户难以察觉。同样,基于视觉语言模型(VLMs)的移动代理在检测隐蔽交互或完成任务动作时面临挑战。为此,我们提出GhostUI,一个专为检测移动应用中隐藏交互而设计的新数据集。该数据集包含操作前后的截图、简化视图层级、手势元数据及任务描述,帮助VLM更好地识别隐藏手势并预测交互后状态。对VLM的定量评估显示,经GhostUI微调的模型在预测隐藏交互和推断交互后屏幕方面表现优于基线模型,凸显了GhostUI作为推进移动端任务自动化基础的潜力。

原文摘要 · Abstract (English)

Modern mobile applications rely on hidden interactions--gestures without visual cues like long presses and swipes--to provide functionality without cluttering interfaces. While experienced users may discover these interactions through prior use or onboarding tutorials, their implicit nature makes them difficult for most users to uncover. Similarly, mobile agents--systems designed to automate tasks on mobile user interfaces, powered by vision language models (VLMs)--struggle to detect veiled interactions or determine actions for completing tasks. To address this challenge, we present GhostUI, a new dataset designed to enable the detection of hidden interactions in mobile applications. GhostUI provides before-and-after screenshots, simplified view hierarchies, gesture metadata, and task descriptions, allowing VLMs to better recognize concealed gestures and anticipate post-interaction states. Quantitative evaluations with VLMs show that models fine-tuned on GhostUI outperform baseline VLMs, particularly in predicting hidden interactions and inferring post-interaction screens, underscoring GhostUI's potential as a foundation for advancing mobile task automation.

移动自动化隐藏交互视觉语言模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。