通过扰动测试揭示多模态智能体的记忆与推理真相
Agent-ScanKit: Unraveling Memory and Reasoning of Multimodal Agents via Sensitivity Perturbations
- 设计三类可控扰动,分别探测视觉、文本和结构信息的作用
- 在18个智能体上发现多数依赖记忆而非系统推理
- 适合关注智能体可靠性与泛化能力的研究者阅读
尽管近期已提出多种策略以提升多模态智能体在图形用户界面(GUI)中的自主交互能力,但在复杂或域外任务下其可靠性仍受限。这引发根本性问题:现有智能体是否在虚假推理?本文提出 extbf{Agent-ScanKit},一种系统性探针框架,通过受控扰动揭示多模态智能体的记忆与推理能力。具体设计三种正交探针范式:视觉引导、文本引导和结构引导,均无需访问模型内部即可量化记忆与推理的贡献。在包含18个多模态智能体的五个公开GUI基准测试中,结果表明机械记忆常压倒系统推理,多数模型主要作为训练对齐知识的检索器,泛化能力有限。研究强调真实场景中构建强推理能力的必要性,为开发可靠多模态智能体提供关键洞见。
原文摘要 · Abstract (English)
Although numerous strategies have recently been proposed to enhance the autonomous interaction capabilities of multimodal agents in graphical user interface (GUI), their reliability remains limited when faced with complex or out-of-domain tasks. This raises a fundamental question: Are existing multimodal agents reasoning spuriously? In this paper, we propose \textbf{Agent-ScanKit}, a systematic probing framework to unravel the memory and reasoning capabilities of multimodal agents under controlled perturbations. Specifically, we introduce three orthogonal probing paradigms: visual-guided, text-guided, and structure-guided, each designed to quantify the contributions of memorization and reasoning without requiring access to model internals. In five publicly available GUI benchmarks involving 18 multimodal agents, the results demonstrate that mechanical memorization often outweighs systematic reasoning. Most of the models function predominantly as retrievers of training-aligned knowledge, exhibiting limited generalization. Our findings underscore the necessity of robust reasoning modeling for multimodal agents in real-world scenarios, offering valuable insights toward the development of reliable multimodal agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。