首个评估多模态智能体在图形界面中可信度的综合框架。
MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments
- 构建四维可信度评估体系:真实性、可控性、安全性和隐私性。
- 在34个高风险交互任务中发现多模态智能体存在累积性风险。
- 适合关注AI代理安全与可信赖性的研究人员和开发者。
多模态大模型智能体(MLA)通过融合视觉、语言、动作与动态环境,实现了从网页自动化到移动系统的自主操作能力。然而,其可执行输出带来的可信度挑战远超传统语言模型,可能直接改变数字状态并引发不可逆现实后果。现有基准无法有效应对MLA的可操作性输出、长周期不确定性及多模态攻击向量等独特问题。本文提出MLA-Trust,首个统一且全面的评估框架,涵盖真实性、可控性、安全性和隐私性四个维度。采用真实网站与移动应用作为测试环境,设计34个高风险交互任务,构建丰富评估数据集。大规模实验涉及13个先进MLA,揭示了多模态交互场景下前所未见的可信度漏洞。例如,专有与开源的GUI交互式MLA比静态多模态大模型带来更严重的可信度风险,尤其在高风险领域;从静态多模态大模型转向交互式智能体显著降低可信度,使多步交互中可能生成有害内容,而独立的大模型通常可避免此类问题;多步执行虽提升适应性,却伴随潜在非线性风险积累,绕过现有防护机制,导致不可预测的衍生风险。此外,我们提供一个可扩展工具箱,支持在多种交互环境中持续评估可信度。
原文摘要 · Abstract (English)
The emergence of multimodal LLM-based agents (MLAs) has transformed interaction paradigms by seamlessly integrating vision, language, action and dynamic environments, enabling unprecedented autonomous capabilities across GUI applications ranging from web automation to mobile systems. However, MLAs introduce critical trustworthiness challenges that extend far beyond traditional language models' limitations, as they can directly modify digital states and trigger irreversible real-world consequences. Existing benchmarks inadequately tackle these unique challenges posed by MLAs' actionable outputs, long-horizon uncertainty and multimodal attack vectors. In this paper, we introduce MLA-Trust, the first comprehensive and unified framework that evaluates the MLA trustworthiness across four principled dimensions: truthfulness, controllability, safety and privacy. We utilize websites and mobile applications as realistic testbeds, designing 34 high-risk interactive tasks and curating rich evaluation datasets. Large-scale experiments involving 13 state-of-the-art agents reveal previously unexplored trustworthiness vulnerabilities unique to multimodal interactive scenarios. For instance, proprietary and open-source GUI-interacting MLAs pose more severe trustworthiness risks than static MLLMs, particularly in high-stakes domains; the transition from static MLLMs into interactive MLAs considerably compromises trustworthiness, enabling harmful content generation in multi-step interactions that standalone MLLMs would typically prevent; multi-step execution, while enhancing the adaptability of MLAs, involves latent nonlinear risk accumulation across successive interactions, circumventing existing safeguards and resulting in unpredictable derived risks. Moreover, we present an extensible toolbox to facilitate continuous evaluation of MLA trustworthiness across diverse interactive environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。