提出检测并纠正计算机使用智能体错误行为的新方法
When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use Agents
- 设计DeAction系统,通过结构化反馈实时检测修正异常操作
- 在真实数据集上提升F1分数超15%,对抗攻击成功率降低90%以上
- 适用于安全要求高的智能体应用,如自动化办公与敏感操作
过去一年中,计算机使用智能体(CUAs)取得显著进展,但仍频繁产生偏离用户原始意图的错误行为。这些错误行为可能源于外部攻击(如间接提示注入)或内部局限(如推理错误),不仅带来安全风险,还降低任务效率与可靠性。本文首次系统定义并研究了CUA中的误动作检测问题,涵盖外部诱导与内部引发的两类误动作。我们识别出实际部署中三种常见场景,构建了包含人类标注动作级对齐标签的真实轨迹基准集MisActBench。进一步提出DeAction,一种实用且通用的防护机制,在动作执行前检测并基于结构化反馈迭代修正误动作。在离线与在线评估中,DeAction均显著优于现有基线,仅引入适度延迟:(1) 在MisActBench上F1分数绝对提升超过15%;(2) 在在线对抗环境下,攻击成功率降低90%以上,同时在正常环境中保持或提升任务成功率。
原文摘要 · Abstract (English)
Computer-use agents (CUAs) have made tremendous progress in the past year, yet they still frequently produce misaligned actions that deviate from the user's original intent. Such misaligned actions may arise from external attacks (e.g., indirect prompt injection) or from internal limitations (e.g., erroneous reasoning). They not only expose CUAs to safety risks, but also degrade task efficiency and reliability. This work makes the first effort to define and study misaligned action detection in CUAs, with comprehensive coverage of both externally induced and internally arising misaligned actions. We further identify three common categories in real-world CUA deployment and construct MisActBench, a benchmark of realistic trajectories with human-annotated, action-level alignment labels. Moreover, we propose DeAction, a practical and universal guardrail that detects misaligned actions before execution and iteratively corrects them through structured feedback. DeAction outperforms all existing baselines across offline and online evaluations with moderate latency overhead: (1) On MisActBench, it outperforms baselines by over 15% absolute in F1 score; (2) In online evaluation, it reduces attack success rate by over 90% under adversarial settings while preserving or even improving task success rate in benign environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。