用自动化方法挖掘电脑代理在正常输入下的严重意外行为
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
- 通过反馈迭代扰动指令,自动触发代理的异常行为
- 从多个前沿代理中发现数百种有害意外行为
- 揭示了不同代理对异常行为的普遍脆弱性
尽管计算机使用代理(CUAs)在自动化复杂操作系统工作流方面潜力巨大,但在良性输入环境下仍可能表现出偏离预期的危险意外行为。然而,此类风险目前多为零散案例,缺乏系统性表征和主动发现长尾意外行为的自动化方法。为此,本文提出首个针对无意中CUA行为的概念与方法框架,定义其关键特征,实现自动诱发并分析这些行为如何由良性输入引发。我们设计了AutoElicit:一种基于执行反馈迭代扰动良性指令的智能体框架,在保持扰动真实且无害的前提下,成功诱发严重危害。利用该框架,我们在Claude 4.5 Haiku、Claude 4.5 Opus及Operator等前沿代理中发现了数百种有害意外行为。进一步评估人类验证的成功扰动迁移性,揭示多种前沿代理普遍存在对意外行为的易感性。本研究为在真实计算机使用场景下系统分析意外行为奠定了基础。
原文摘要 · Abstract (English)
Although computer-use agents (CUAs) hold significant potential to automate increasingly complex OS workflows, they can demonstrate unsafe unintended behaviors that deviate from expected outcomes even under benign input contexts. However, exploration of this risk remains largely anecdotal, lacking concrete characterization and automated methods to proactively surface long-tail unintended behaviors under realistic CUA scenarios. To fill this gap, we introduce the first conceptual and methodological framework for unintended CUA behaviors, by defining their key characteristics, automatically eliciting them, and analyzing how they arise from benign inputs. We propose AutoElicit: an agentic framework that iteratively perturbs benign instructions using CUA execution feedback, and elicits severe harms while keeping perturbations realistic and benign. Using AutoElicit, we surface hundreds of harmful unintended behaviors from state-of-the-art CUAs such as Claude 4.5 Haiku, Claude 4.5 Opus, and Operator. We further evaluate the transferability of human-verified successful perturbations, identifying persistent susceptibility to unintended behaviors across various other frontier CUAs. This work establishes a foundation for systematically analyzing unintended behaviors in realistic computer-use settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。