通过精确评估定位偏差,提升GUI智能体的可靠性与可解释性。
Understanding GUI Agent Localization Biases through Logit Sharpness
- 按预测类型细分为四类,揭示传统准确率忽略的错误模式。
- 提出峰度得分(PSS)量化模型不确定性,关联语义连续性与概率分布。
- 无需训练的上下文自适应裁剪,显著改善定位精度,适合系统开发者使用。
多模态大语言模型(MLLMs)使GUI智能体能够通过将语言映射到空间动作来与操作系统交互。尽管性能令人期待,这些模型常出现幻觉——系统性定位错误,影响可靠性。本文提出一个细粒度评估框架,将模型预测分为四类,揭示了传统准确率指标无法捕捉的细微失败模式。为更精准量化模型不确定性,引入峰值尖度得分(Peak Sharpness Score, PSS),评估坐标预测中语义连续性与logits分布的一致性。基于此洞察,进一步提出无须训练的上下文自适应裁剪(Context-Aware Cropping)技术,通过动态优化输入上下文提升模型表现。大量实验表明,该框架与方法提供可操作洞见,增强GUI智能体行为的可解释性与鲁棒性。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have enabled GUI agents to interact with operating systems by grounding language into spatial actions. Despite their promising performance, these models frequently exhibit hallucinations-systematic localization errors that compromise reliability. We propose a fine-grained evaluation framework that categorizes model predictions into four distinct types, revealing nuanced failure modes beyond traditional accuracy metrics. To better quantify model uncertainty, we introduce the Peak Sharpness Score (PSS), a metric that evaluates the alignment between semantic continuity and logits distribution in coordinate prediction. Building on this insight, we further propose Context-Aware Cropping, a training-free technique that improves model performance by adaptively refining input context. Extensive experiments demonstrate that our framework and methods provide actionable insights and enhance the interpretability and robustness of GUI agent behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。