arXiv:2510.02204cs.CL2025-10被引 6

诊断视觉语言模型在手机任务中的推理与执行脱节问题

Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents

  • 提出新评估框架,检验推理是否真实指向正确操作
  • 发现大模型仍存在显著执行错误,且比推理错误更常见
  • 适合关注智能助手可信性与安全性的研究者

基于视觉语言模型的手机使用代理在理解自然语言指令并生成相应操作方面展现出巨大潜力。尽管链式思维(CoT)推理常被证明能提升执行准确率,但现有评估多聚焦执行结果,忽略推理过程与真实操作的一致性。这种忽视导致无法检测推理-执行差距,进而引发过度信任:用户可能因看似合理的推理而授权有害操作,造成经济损失或信任危机。本文提出新评估框架,核心为真值对齐(GTA),衡量CoT所隐含的操作是否与真实操作一致。结合标准精确匹配(EM)指标,联合评估推理与执行准确性。实验显示,两类差距普遍存在:执行差距(EG)指推理正确但执行失败;推理差距(RG)指执行成功但推理与实际操作矛盾。在多种手机交互任务中,执行差距占比更高。即使最大模型,执行差距依然显著。分析表明该框架可有效揭示先进模型中的系统性模式,为构建更可信的手机代理提供诊断支持。

原文摘要 · Abstract (English)

Mobile-use agents powered by vision-language models (VLMs) have shown great potential in interpreting natural language instructions and generating corresponding actions based on mobile graphical user interface. Recent studies suggest that incorporating chain-of-thought (CoT) reasoning tends to improve the execution accuracy. However, existing evaluations emphasize execution accuracy while neglecting whether CoT reasoning aligns with ground-truth actions. This oversight fails to assess potential reasoning-execution gaps, which in turn foster over-trust: users relying on seemingly plausible CoTs may unknowingly authorize harmful actions, potentially resulting in financial loss or trust crisis. In this work, we introduce a new evaluation framework to diagnose reasoning-execution gaps. At its core lies Ground-Truth Alignment (GTA), which measures whether the action implied by a CoT matches the ground-truth action. By combining GTA with the standard Exact Match (EM) metric, we jointly assess both the reasoning accuracy and execution accuracy. This joint perspective reveals two types of reasoning-execution gaps: (i) Execution Gap (EG), where the reasoning correctly identifies the correct action but execution fails, and (ii) Reasoning Gap (RG), where execution succeeds but reasoning process conflicts with the actual execution. Experimental results across a wide range of mobile interaction tasks reveal that reasoning-execution gaps are prevalent, with execution gaps occurring more frequently than reasoning gaps. Moreover, while scaling up model size reduces the overall gap, sizable execution gaps persist even in the largest models. Further analysis shows that our framework reliably reflects systematic EG/RG patterns in state-of-the-art models. These findings offer concrete diagnostics and support the development of more trustworthy mobile-use agents.

视觉语言模型推理验证智能代理可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。