arXiv:2608.01604cs.AIcs.SE2026-08

在办公任务上后训练,能提升代码生成能力。

Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross-Domain Transfer

论文配图:Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross-Domain Transfer
图 1 · 摘自论文原文
  • 通过办公场景长程任务训练模型,强化目标导向执行行为
  • 代码任务通过率提升5.8分,四类核心行为均获改善
  • 效果跨域迁移,适合关注模型泛化与行为建模的研究者

长周期任务要求智能体在嵌套和分支工作中持续保持一致的目标与状态。我们提出目标导向执行(GDE)能力,包含四项行为:选择目标、构建相关状态、维持高层目标一致性、环境验证完成度。假设长周期后训练可增强跨领域这些行为。我们在363个来自办公流程的长周期多工具任务上对Qwen3.5-122B-A10B进行后训练,该数据集不含软件工程任务,但模型在SWE-Bench Pro上的pass@1提升5.8分。轨迹匹配分析显示,办公室工作与代码仓库中四类GDE行为均有提升。整体统计显示信息获取、实现与验证环节均出现相关改进。结果支持一种行为解释:长周期后训练改变了模型在任务间组织与应用知识的方式,且效果超越训练领域。

原文摘要 · Abstract (English)

Long-horizon tasks require agents to maintain coherent state and goals across nested and branching work. We call this capability goal-directed execution (GDE): the repeated application of four behaviors, namely selecting goals, constructing task-relevant state, maintaining fidelity to higher-level objectives, and verifying completion against the environment. We hypothesize that long-horizon post-training strengthens these behaviors across domains. We test this by post-training Qwen3.5-122B-A10B on 363 Long-Horizon Multi-Tool Agent (LHMTA) tasks drawn from office workflows. The collection contained no software-engineering tasks, yet the model's pass@1 improved by 5.8 points on SWE-Bench Pro. Matched trajectory analysis shows gains in all four GDE behaviors in both office workflows and software repositories. Aggregate SWE-Bench Pro statistics showed related changes in information gathering, implementation, and verification. Together, the results support a behavioral interpretation in which long-horizon post-training changed how the model organized and applied knowledge across tasks, with effects extending beyond the training domain.

长程任务跨域迁移目标导向代码生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。