首个评估多模态模型操控电脑风险的基准,揭示真实场景中的安全漏洞。
RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
- 构建涵盖492个任务的真实交互环境,覆盖网页、社交、办公等多类应用
- 发现当前计算机代理在真实操作中存在显著安全风险,风险完成率较高
- 区分用户与环境来源风险,为可信智能代理开发提供实证依据
随着多模态大语言模型(MLLMs)快速发展,它们被越来越多地用作能够完成复杂计算机任务的自主代理。然而,一个紧迫问题浮现:专为对话场景设计和对齐的安全原则能否有效迁移到真实世界计算机操作场景?现有研究在评估基于MLLM的计算机代理安全风险时存在多重局限:或缺乏真实交互环境,或仅聚焦单一或少数风险类型。这些局限忽视了真实环境的复杂性、多样性和变异性,制约了对计算机代理的全面风险评估。为此,我们提出 extbf{RiOSWorld},一个用于评估多模态模型代理在真实世界计算机操作中潜在风险的基准。该基准包含492个涉及网页、社交媒体、多媒体、操作系统、邮件及办公软件等各类应用的危险任务。我们根据风险来源将其分为两类:(i) 用户起源风险和 (ii) 环境风险。评估从两个维度进行:(i) 风险目标意图和 (ii) 风险目标完成度。在 extbf{RiOSWorld}上对多模态代理的大量实验表明,当前计算机代理在真实场景中面临显著安全风险。研究结果凸显了在真实计算机操作中进行安全对齐的必要性与紧迫性,为构建可信计算机代理提供了重要启示。基准已公开:https://yjyddq.github.io/RiOSWorld.github.io/
原文摘要 · Abstract (English)
With the rapid development of multimodal large language models (MLLMs), they are increasingly deployed as autonomous computer-use agents capable of accomplishing complex computer tasks. However, a pressing issue arises: Can the safety risk principles designed and aligned for general MLLMs in dialogue scenarios be effectively transferred to real-world computer-use scenarios? Existing research on evaluating the safety risks of MLLM-based computer-use agents suffers from several limitations: it either lacks realistic interactive environments, or narrowly focuses on one or a few specific risk types. These limitations ignore the complexity, variability, and diversity of real-world environments, thereby restricting comprehensive risk evaluation for computer-use agents. To this end, we introduce \textbf{RiOSWorld}, a benchmark designed to evaluate the potential risks of MLLM-based agents during real-world computer manipulations. Our benchmark includes 492 risky tasks spanning various computer applications, involving web, social media, multimedia, os, email, and office software. We categorize these risks into two major classes based on their risk source: (i) User-originated risks and (ii) Environmental risks. For the evaluation, we evaluate safety risks from two perspectives: (i) Risk goal intention and (ii) Risk goal completion. Extensive experiments with multimodal agents on \textbf{RiOSWorld} demonstrate that current computer-use agents confront significant safety risks in real-world scenarios. Our findings highlight the necessity and urgency of safety alignment for computer-use agents in real-world computer manipulation, providing valuable insights for developing trustworthy computer-use agents. Our benchmark is publicly available at https://yjyddq.github.io/RiOSWorld.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。