无需人工标注,自动适应新电脑环境的持续学习框架
Autonomous Continual Learning for Environment Adaptation of Computer-Use Agents
- 用自动生成任务+反馈循环实现自主持续训练
- 在目标环境中性能提升3%-29%,且不遗忘旧环境
- 适合需要长期适应复杂数字环境的自动化工具
现实世界中的数字环境高度多样且动态变化,导致计算机使用代理(CUAs)频繁遭遇未见环境与分布偏移,因此在该类环境中进行持续学习至关重要。然而,获取高质量且与环境相关的训练数据往往依赖昂贵的人工标注。本文提出ACuRL,一种自主课程强化学习框架,可零人工数据下持续适应特定环境。代理首先探索环境获取初始经验;后续迭代训练中,课程任务生成器结合这些经验与前一阶段反馈,生成适配当前能力的新任务。为提供可靠奖励信号,引入CUAJudge,一种鲁棒的自动评估器,在人类判断上达到93%一致率。实验证明,该方法有效实现环境内与跨环境持续学习,在目标环境上获得3%-29%绝对性能提升,且对其他环境无灾难性遗忘。还能缓解环境变化(如版本更新、平台迁移、分辨率变化)带来的性能下降。进一步分析显示,仅需约20%参数更新,解释了其高效稳健的适应能力。
原文摘要 · Abstract (English)
Real-world digital environments are highly diverse and dynamic. These characteristics cause agents to frequently encounter unseen environments and distribution shifts, making continual learning in such environments essential for computer-use agents (CUAs). However, a key challenge lies in obtaining high-quality and environment-grounded training data without relying on costly human annotation. In this work, we introduce ACuRL, an Autonomous Curriculum Reinforcement Learning framework that continually adapts agents to specific environments with zero human data. The agent first explores an environment to acquire initial experiences. During subsequent iterative training, a curriculum task generator leverages these experiences together with feedback from the previous iteration to synthesize new tasks tailored for the agent's current capabilities. To provide reliable reward signals, we introduce CUAJudge, a robust automatic evaluator for CUAs that achieves 93% agreement with human judgments. Empirically, our method effectively enables both intra-environment and cross-environment continual learning, yielding 3-29% absolute performance gains on the target environments without catastrophic forgetting on others. We also show that it can mitigate performance degradation under environment changes (e.g., version updates, platform migration, and resolution shifts). Further analyses show highly sparse updates (e.g., only 20% parameters), which helps explain the effective and robust adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。