构建首个覆盖七类真实异常的GUI代理评测数据集,揭示现有模型在异常下的性能暴跌
GUI-Robust: A Comprehensive Dataset for Testing GUI Agent Robustness in Real-World Anomalies
- 用RPA工具采集真实用户操作,结合多模态大模型自动生成标注,效率提升19倍以上
- 测试发现主流GUI代理在异常场景下性能显著下降,暴露出鲁棒性短板
- 适合关注GUI智能体可靠性、自动化测试与真实世界部署的研究者
高质量数据集对推动图形用户界面(GUI)智能体研究至关重要。然而,现有数据集多在理想化条件下构建,忽视了实际应用中常见的多样化异常。为此,我们提出GUI-Robust,一个专为全面评估GUI代理鲁棒性而设计的新数据集,明确包含日常交互中常见的七类异常。同时,我们提出一种半自动化数据构建范式:通过RPA工具收集自然用户操作序列,并借助多模态大模型(MLLMs)自动生成对应的操作步骤与任务描述,使标注成本降低超过19倍。最后,我们使用GUI-Robust对当前最先进的GUI代理进行评估,发现其在异常场景下表现显著退化。本工作强调了鲁棒性在GUI代理中的重要性,有望激发更多相关研究。数据集与代码已开源:https://github.com/chessbean1/GUI-Robust。
原文摘要 · Abstract (English)
The development of high-quality datasets is crucial for benchmarking and advancing research in Graphical User Interface (GUI) agents. Despite their importance, existing datasets are often constructed under idealized conditions, overlooking the diverse anomalies frequently encountered in real-world deployments. To address this limitation, we introduce GUI-Robust, a novel dataset designed for comprehensive GUI agent evaluation, explicitly incorporating seven common types of anomalies observed in everyday GUI interactions. Furthermore, we propose a semi-automated dataset construction paradigm that collects user action sequences from natural interactions via RPA tools and then generate corresponding step and task descriptions for these actions with the assistance of MLLMs. This paradigm significantly reduces annotation time cost by a factor of over 19 times. Finally, we assess state-of-the-art GUI agents using the GUI-Robust dataset, revealing their substantial performance degradation in abnormal scenarios. We anticipate that our work will highlight the importance of robustness in GUI agents and inspires more future research in this direction. The dataset and code are available at https://github.com/chessbean1/GUI-Robust..
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。