为视觉语言模型的界面操作代理设计跨模型、跨数据集的不确定性评估基准。
Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets

- 构建27种开源方法与8种闭源方法的跨域评估矩阵,覆盖4个模型和4个数据集。
- 发现同一模型内不确定性排序稳定(斯皮尔曼相关达0.969),但跨模型则大幅下降。
- 闭源模型需在目标环境重排优劣,单纯依赖外部评估不可靠。
计算机使用代理将视觉语言模型(VLM)预测转化为可执行的GUI点击操作,因此可靠的不确定性估计对拒绝决策、校准、误判严重性排序及空间安全区域至关重要。然而,现有后验不确定性量化(UQ)研究分散于孤立的模型与数据集组合,难以判断当代理、基准或可观测界面变化时,UQ排序是否保持稳定。本文提出Argus,一个面向单步可执行GUI定位的跨范式基准:包含4个VLM代理与4个数据集的27种开源方法矩阵,以及3个前沿厂商的8种闭源方法矩阵(无法获取logits、隐藏状态与注意力图)。评估方法涵盖基于logits的分数、采样与一致性度量、隐藏状态与密度估计(马氏距离、SAPLMA)、基于注意力的评分、P(True)与口语化置信度提示、以及分组合取预测。主要发现为选择性迁移:固定模型下UQ排序在不同数据集间稳定,但跨模型类别与可观测界面时性能下降。隐藏状态与密度类方法在开源中最为稳定;而CoCoA-1MCA、Focus、采样类分数与口语化自评在特定场景表现最优。模型内排序迁移性强(斯皮尔曼ρ最高达0.969),但跨层级迁移至闭源厂商平均仅+0.08,表明闭源UQ应针对目标环境重新排序而非外推。合取点击区域显示,仅分数层面区分不足:经校准后局部加权圆盘半径缩小40-60%,但校准测试或界面不匹配下覆盖率下降。我们开放每项记录、校准/测试划分、UQ分数与分析脚本,支持针对具体场景选择合适的不确定性量化策略。
原文摘要 · Abstract (English)
Computer-use agents turn vision-language model (VLM) predictions into executable GUI clicks, so reliable uncertainty estimates are essential for rejection, calibration, miss-severity ranking, and spatial safety regions. Yet evidence on post-hoc uncertainty quantification (UQ) for these agents is fragmented across isolated model and dataset pairs, leaving it unclear whether UQ rankings stay stable when the agent, benchmark, or observable interface changes. We present Argus, a cross-regime benchmark for post-hoc UQ in single-step executable GUI grounding: a 27-method open-weight matrix over 4 VLM agents and 4 datasets, plus an 8-method closed-source matrix across 3 frontier vendors where logits, hidden states, and attention maps are unavailable. Evaluated methods span logit-based scores, sampling and consistency measures, hidden-state and density estimators (Mahalanobis, SAPLMA), attention-based scores, P(True) and verbalised-confidence prompting, and split-conformal prediction. The main finding is selective transfer: UQ rankings are stable across datasets for a fixed model, but degrade across model classes and observable interfaces. Hidden-state and density methods are the most stable open-weight family, while CoCoA-1MCA, Focus, sampling-based scores, and verbalised self-assessment win in specific regimes. Within-model ranking transfer is strong (Spearman rho up to 0.969), but cross-tier transfer to closed-source vendors averages only +0.08, so closed-source UQ should be reranked on the target rather than extrapolated. Conformal click regions show score-level discrimination is not enough for deployment: locally weighted disks shrink radii by 40-60% when the plug-in UQ is calibrated, but coverage degrades under calibration-test or interface mismatch. We release per-item records, calibration/test splits, UQ scores, and analysis scripts for regime-aware UQ selection in GUI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。