构建多语言GUI智能体评测基准,提升非英语场景下的感知与推理能力。
MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents
- 设计六语言对齐的细粒度评测框架,精准定位任务失败原因。
- 发现非英语环境下推理能力普遍落后,尤其在复杂任务中差距显著。
- 提出GUI-XLI干预方法,通过隐状态对齐提升非英语表现,平均增益6.5%。
大型视觉语言模型在多语言图形用户界面(GUI)智能体方面展现出强大潜力,但现有评测基准存在两大缺陷:一是缺乏对感知与推理(P&R)能力的细粒度诊断,难以定位任务失败根源;二是缺少严格对齐的跨语言评估环境,导致语言影响无法独立分析。为此,我们提出多语言P&R GUI基准(MPR-GUI-Bench),涵盖六种语言和八项细粒度P&R任务,实现跨语言环境严格对齐。实验揭示英语与非英语设置间存在持续的P&R差距,尤其在推理密集型任务上。为利用英语优异的P&R能力弥补跨语言鸿沟,我们识别出对语言敏感的模型层,并提出GUI-XLI方法,在推理时将非英语隐藏状态对齐至对应英语表示。结果表明,GUI-XLI有效缩小跨语言差距,非英语场景下平均性能提升6.5%。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have shown strong potential as multilingual Graphical User Interface (GUI) agents, as evidenced by existing GUI benchmarks. However, these benchmarks exhibit two primary limitations: (1) although Perception and Reasoning (P&R) capabilities are fundamental for GUI agents, current benchmarks lack fine-grained diagnostics to identify which specific capabilities lead to task failures, hindering targeted improvements; (2) existing benchmarks fail to provide a strictly aligned cross-lingual evaluation environment, introducing confounding factors that prevent isolating the language impact on GUI agent performance. To address these issues, we propose the Multilingual P&R GUI Benchmark (MPR-GUI-Bench), featuring strictly aligned environments across six languages and eight fine-grained P&R tasks. Our benchmark reveals consistent P&R gaps between English and non-English settings, particularly on reasoning-intensive tasks. To leverage the superior English P&R capabilities for bridging cross-lingual gaps, we identify layers sensitive to language and propose GUI-XLI, a GUI Cross-Lingual Intervention method that aligns non-English hidden states with their English counterparts at these layers during inference. Experiments show that GUI-XLI effectively reduces the cross-lingual gaps, with an average gain of 6.5% in non-English settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。