让小模型通过识别弱点实现高效领域专化,效果远超盲目训练。
Learn from Weaknesses: Automated Domain Specialization for Small Computer-Use Agents

- 用强模型发现弱模型在目标领域的短板,针对性生成训练任务。
- 在8个领域上分别提升11.6%和11.1%,显著优于基线模型。
- 强调学生感知的数据与训练机制,适合资源有限的实用型智能体开发。
计算机使用智能体(CUAs)近年取得显著进展,但为每个软件领域部署独立大模型成本高昂。小型开源智能体更适合作为专化目标,但其能力仍较弱且存在不均衡的领域失败。直接合成大规模目标域训练数据虽可行,但效果仅小幅提升。为此,本文提出无标注专化框架LearnWeak:利用更强参考智能体识别学生模型在目标域的薄弱环节,自动生成针对性任务并构建监督信号。该方法引入误差感知的专化目标,分离规划与执行错误,实现更精准的行为优化。在OSWorld上,LearnWeak在八个领域中相较EvoCUA-8B和OpenCUA-7B分别提升11.6%和11.1%。验证还表明,学生感知的数据生成与训练策略优于现有自主轨迹生成与训练基线。研究强调了在数据合成与训练中融入学生意识的重要性,为小规模计算机使用智能体在多样化领域的高效专化提供了更系统、高效的新路径。
原文摘要 · Abstract (English)
Computer-use agents (CUAs) have recently made substantial progress, but deploying a separate large expert for each software domain remains expensive. Small open computer-use agents are more practical specialization targets, but they remain substantially weaker and exhibit uneven domain-specific failures. A straightforward remedy is to synthesize large-scale training data for the target domain, yet we find that this naive approach yields only marginal improvements. Building on this observation, we introduce LearnWeak, an annotation-free specialization framework for small computer-use agents that uses a stronger reference agent to identify the student's weaknesses in the target domain, synthesize targeted tasks, and construct supervision automatically. LearnWeak further introduces an error-aware specialization objective that disentangles planning and execution errors, enabling more behaviorally precise updates than broad uniform supervision. On OSWorld, LearnWeak achieves average gains of 11.6 and 11.1 percentage points over EvoCUA-8B and OpenCUA-7B, respectively, across eight domains. We also validate that our student-aware dataset generation and training approaches outperform existing autonomous trajectory generation and training baselines. Our work highlights the importance of student awareness in both data synthesis and agent training, pointing toward a more principled and efficient path for specializing small computer-use agents in diverse domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。