arXiv:2604.18847cs.AIcs.CL2026-04

让智能助手在误操作后能按人意愿高效修复,避免损失。

Human-Guided Harm Recovery for Computer Use Agents

论文配图:Human-Guided Harm Recovery for Computer Use Agents
图 1 · 摘自论文原文
  • 通过用户研究提取修复偏好,构建自然语言评价标准
  • 用1130组对比判断发现:短期精准修复更受青睐
  • 新奖励模型让助手在50个任务中修复效果优于基线

随着大模型代理在真实计算机系统上执行操作,我们不仅需要大规模预防有害行为,还需在预防失败时有效补救。本文提出‘伤害恢复’这一新问题:如何在符合人类偏好的前提下,最优地将代理从有害状态引导回安全状态。通过一项形成性用户研究,识别出关键的恢复维度并生成自然语言评价标准。基于1,130组成对判断的数据集揭示,属性重要性具有上下文依赖性,例如人们更倾向于实用、精准的短期策略而非全面长期方案。我们将这些发现融入奖励模型,在测试时对代理生成的多个恢复方案进行重排序。为系统评估恢复能力,我们引入BackBench基准,包含50个计算机使用任务,用于检验代理在有害状态下的恢复能力。人工评估表明,采用该奖励模型的代理生成的恢复轨迹质量高于基础代理和基于规则的支架。这些成果为新一代代理安全方法奠定基础——既防患于未然,也善后处理已发生的损害。

原文摘要 · Abstract (English)

As LM agents gain the ability to execute actions on real computer systems, we need ways to not only prevent harmful actions at scale but also effectively remediate harm when prevention fails. We formalize a solution to this neglected challenge in post-execution safeguards as harm recovery: the problem of optimally steering an agent from a harmful state back to a safe one in alignment with human preferences. We ground preference-aligned recovery through a formative user study that identifies valued recovery dimensions and produces a natural language rubric. Our dataset of 1,130 pairwise judgments reveals context-dependent shifts in attribute importance, such as preferences for pragmatic, targeted strategies over comprehensive long-term approaches. We operationalize these learned insights in a reward model, re-ranking multiple candidate recovery plans generated by an agent scaffold at test time. To evaluate recovery capabilities systematically, we introduce BackBench, a benchmark of 50 computer-use tasks that test an agent's ability to recover from harmful states. Human evaluation shows our reward model scaffold yields higher-quality recovery trajectories than base agents and rubric-based scaffolds. Together, these contributions lay the foundation for a new class of agent safety methods -- ones that confront harm not only by preventing it, but by navigating its aftermath with alignment and intent.

智能代理安全修复人机对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。