arXiv:2608.11434cs.AIcs.CL2026-08

测试大模型当裁判评估手机智能体表现,发现简单方法反而更可靠。

Benchmarking LLM Judges for Mobile Agent Evaluation

论文配图:Benchmarking LLM Judges for Mobile Agent Evaluation
图 1 · 摘自论文原文
  • 用931条真人标注轨迹测试6种大模型裁判方法。
  • 简单截图采样法常优于复杂设计,模型底座决定裁判效果。
  • 不同大模型裁判有相反错误模式,影响实际应用选择。

移动智能体评测日益依赖大模型作为裁判判断任务完成情况,但其在移动智能体轨迹上的可靠性尚未充分研究。本文提出MobileJudgeBench,一个系统性评估大模型裁判方法在移动智能体轨迹上表现的基准。该基准包含931条人类标注轨迹,覆盖6个移动智能体基准、4种智能体模型和68个应用。我们在多个大模型后端上评估了6种裁判方法(五种来自SPA-Bench、A3两种模式、AndroidArena、AgentRewardBench,以及我们设计的一个简单基线)。实验揭示三个关键发现:第一,简单的基于采样截图的基线裁判方法在多数情况下与专门设计的方法相当甚至更优,表明更复杂的裁判流程并不总能提升质量;在表现相近的方法中,大模型底座是主要决定因素。第二,基准质量指标可有效预测真实裁判效用:它们与智能体排名保真度以及作为奖励信号用于在线强化学习时的下游性能高度相关。第三,对两种大模型后端的失败分析揭示了定性相反的失败特征——一种保守,一种宽松,分别对应于底座模型的精确率-召回率特性。

原文摘要 · Abstract (English)

Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamined. We introduce MobileJudgeBench, a benchmark for systematically evaluating LLM-as-judge methods on mobile agent trajectories. Our benchmark comprises 931 human-annotated trajectories spanning 6 mobile agent benchmarks, 4 agent models, and 68 apps. We evaluate 6 judge methods (five adapted from SPA-Bench, A3 with two modes, AndroidArena, and AgentRewardBench, plus a simple baseline we design) across multiple LLM backends. Our experiments reveal three key findings. First, a simple baseline judge with sampled screenshots is competitive with, and often exceeds, purpose-built methods, indicating that more elaborate judge pipelines do not consistently improve judge quality; among competitive methods, the LLM backbone is the primary driver. Second, benchmark quality metrics reliably predict real-world judge utility: they correlate with both agent ranking fidelity for evaluation and downstream performance when judges serve as reward signals for on-policy reinforcement learning. Third, failure analysis across two LLM backends uncovers qualitatively opposite failure profiles, one conservative and the other permissive, linked to the backbone's precision-recall characteristics.

大模型评测智能体评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。