测试主流GUI定位模型在攻击下的鲁棒性,发现其极易受干扰。
On the Robustness of GUI Grounding Models Against Image Attacks
- 在多种界面环境下测试模型对噪声和对抗攻击的响应
- 模型在对抗攻击下准确率显著下降,低分辨率时更脆弱
- 研究结果为提升智能交互系统可靠性提供重要参考
图形用户界面(GUI)定位模型对于智能体理解与操作复杂视觉界面至关重要。然而,这些模型在真实场景中易受自然噪声和对抗扰动影响,其鲁棒性尚未得到充分研究。本文系统评估了当前主流的GUI定位模型(如UGround)在三种条件下的表现:自然噪声、非目标对抗攻击和目标对抗攻击。实验覆盖移动端、桌面端和网页端等多种界面环境,结果表明,GUI定位模型对对抗扰动和低分辨率条件高度敏感。该研究揭示了现有模型的关键弱点,并为未来提升其实用场景鲁棒性提供了基准支持。代码已开源:https://github.com/ZZZhr-1/Robust_GUI_Grounding。
原文摘要 · Abstract (English)
Graphical User Interface (GUI) grounding models are crucial for enabling intelligent agents to understand and interact with complex visual interfaces. However, these models face significant robustness challenges in real-world scenarios due to natural noise and adversarial perturbations, and their robustness remains underexplored. In this study, we systematically evaluate the robustness of state-of-the-art GUI grounding models, such as UGround, under three conditions: natural noise, untargeted adversarial attacks, and targeted adversarial attacks. Our experiments, which were conducted across a wide range of GUI environments, including mobile, desktop, and web interfaces, have clearly demonstrated that GUI grounding models exhibit a high degree of sensitivity to adversarial perturbations and low-resolution conditions. These findings provide valuable insights into the vulnerabilities of GUI grounding models and establish a strong benchmark for future research aimed at enhancing their robustness in practical applications. Our code is available at https://github.com/ZZZhr-1/Robust_GUI_Grounding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。