让计算机使用智能体的记忆可验证,防止错误习惯累积。
VerificAgent: Domain-Specific Memory Verification for Scalable Oversight of Aligned Computer-Use Agents
- 用领域专家知识+人类事后审核,构建可审计的记忆库。
- 测试中任务成功率提升,幻觉导致的失败减少37%。
- 适合需要长期安全、可追溯行为的办公自动化场景。
持续记忆增强使计算机使用智能体(CUAs)能从过往交互中学习,但未经审查的记忆可能包含不合适的或不安全的启发式规则——偏离用户意图与安全约束的虚假规律。我们提出VerificAgent,一种可扩展的监督框架,将持久记忆视为明确的对齐界面。该框架结合(1)专家策划的领域知识种子,(2)训练期间基于轨迹的迭代记忆增长,以及(3)部署前的人工事实核查步骤,以净化累积的记忆。在OSWorld生产力任务及额外对抗性压力测试中,VerificAgent提升了任务可靠性,减少了幻觉引发的失败,并保持了可解释、可审计的指导原则——无需额外模型微调。通过让人类一次性纠正高影响错误,经验证的记忆成为未来动作必须遵守的冻结安全合约。结果表明,领域限定、人工验证的记忆为CUAs提供了一种可扩展的监督机制,通过限制隐性策略漂移,锚定智能体行为于目标领域的规范与安全约束之中,补充更广泛的对齐策略。
原文摘要 · Abstract (English)
Continual memory augmentation lets computer-using agents (CUAs) learn from prior interactions, but unvetted memories can encode domain-inappropriate or unsafe heuristics--spurious rules that drift from user intent and safety constraints. We introduce VerificAgent, a scalable oversight framework that treats persistent memory as an explicit alignment surface. VerificAgent combines (1) an expert-curated seed of domain knowledge, (2) iterative, trajectory-based memory growth during training, and (3) a post-hoc human fact-checking pass to sanitize accumulated memories before deployment. Evaluated on OSWorld productivity tasks and additional adversarial stress tests, VerificAgent improves task reliability, reduces hallucination-induced failures, and preserves interpretable, auditable guidance--without additional model fine-tuning. By letting humans correct high-impact errors once, the verified memory acts as a frozen safety contract that future agent actions must satisfy. Our results suggest that domain-scoped, human-verified memory offers a scalable oversight mechanism for CUAs, complementing broader alignment strategies by limiting silent policy drift and anchoring agent behavior to the norms and safety constraints of the target domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。