arXiv:2608.06352cs.LGcs.CL2026-08

用对抗校准生成更适配智能体学习的挑战性任务

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

论文配图:CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks
图 1 · 摘自论文原文
  • 通过对抗式求解器校准动态优化任务难度
  • 生成5431个校准任务,提升模型性能最高达30.04点
  • 适合构建可迁移的智能体训练数据集

训练终端智能体需要可执行且可验证的任务,不仅要求可解,还需对学习具有适当挑战性。可执行验证仅说明可行性,无法反映任务在特定求解器设置下的表现。本文提出CalibForge,一个基于已验证求解器行为的自主终端任务生成系统,通过对抗求解器校准来修正候选任务。多求解器校准针对异构求解器池中的分歧,对比校准则聚焦于强通过/弱失败关系;两者均将学习有效区锚定在可求解性基础上。利用CalibForge,我们构建了5,431个校准后的终端任务。消融实验表明,两种策略相比单纯人工设计或单求解器反馈,能提供更有效的监督。在全量任务上训练的模型在Terminal-Bench 2.0上分别达到32.58%和47.57%的准确率,相较基线模型最大提升达24.71个百分点(Terminal-Bench 2.0)、27.68点(SWE-bench Pro)和30.04点(Doc2Repo)。结果支持求解器相对可学习性作为构建高效、可迁移智能体训练数据的可行目标。

原文摘要 · Abstract (English)

Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a given solver setting. In this paper, we present CalibForge, an autonomous terminal-task synthesis system that uses verified solver behavior to revise candidate tasks through adversarial solver calibration. Multi-solver calibration targets disagreement within a heterogeneous solver pool, whereas contrastive solver calibration targets a designated strong-pass/weak-fail relation; both operationalize a solver-relative learnable zone anchored in demonstrated solvability. Using CalibForge, we construct 5,431 calibrated terminal tasks. Our ablations show that both strategies yield more effective supervision than authoring and validation alone or ordinary single-solver feedback. Models trained on the full collection achieve 32.58% and 47.57% on Terminal-Bench 2.0. The largest improvements over the corresponding base model reach 24.71 percentage points on Terminal-Bench 2.0, 27.68 points on SWE-bench Pro, and 30.04 points on Doc2Repo. Together, these results support solver-relative learnability as a practical target for constructing effective and transferable agent training data.

智能体训练任务生成对抗校准可迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。