arXiv:2608.18324cs.AI2026-08

用验证器筛选的执行记录训练模型,实现高效可靠的流程修复。

Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair

  • 用验证器选中的执行记录作为监督信号,训练无思考能力的模型。
  • 在80个新任务中,接受计划数从1提升至57,且无失败案例。
  • 仅需原推理时间的1/56,适合对效率要求高的实际部署场景。

机器可验证的工作流生成治理记录,包含任务合约、模型尝试、验证决策、接受输出和目标来源。我们测试这些记录能否监督有限能力的模型,将偶尔或昂贵的能力转化为可靠的一次性执行。在全新的结构不匹配的PlanBench重规划案例中,Qwen3-14B的思考模式生成了24个被独立编写的VAL验证器接受的计划。这些计划用于训练同一检查点的非思考执行,无需真值目标或更强教师模型。在80个未开启的案例中,被VAL接受的计划数从1增至57,其中56例有提升,零例退化;非思考模式达到30个。该适配器在所有案例中保持模式有效性,平均延迟仅为思考模式的1/56。独立的配对接口修复门未通过。对照消融实验固定源案例、52候选池、24目标数量、模型、配方与种子,仅改变目标选择方式。在160个新案例中,基础模型、模式选择、模型自选与VAL选择分别获得1、55、69、102个接受计划。VAL显著优于自选,净增33(p=0.0000019647),且在不同难度层级均有效。因此,在此范围内,独立语义选择是核心支撑。另一项使用更强教师模型(Phi)的补充实验,将基础Phi-4的接受计划数从2提升至51,模式有效性输出从35升至80。早期合成实验验证了可教学性、累积学习、构建鲁棒性和停止边界。结果支持验证器选择的监督机制适用于有限、机器可检的能力,而非任意规划、企业合规或无限制自我改进。

原文摘要 · Abstract (English)

Machine-verifiable workflows produce governance records linking a task contract, model attempt, verifier decision, accepted output, and target origin. We test whether these records can supervise bounded models, consolidating occasional or expensive capability into reliable one-shot execution. On fresh, structure-disjoint PlanBench replanning cases, Qwen3-14B thinking generated 24 plans admitted by the independently authored VAL verifier. Those plans trained the same checkpoint for non-thinking execution, without oracle targets or a stronger teacher. On 80 unopened cases, VAL-accepted plans increased from 1 to 57, with 56 paired gains and zero regressions; thinking reached 30. The adapter was schema-valid on all cases and used approximately 1/56 of thinking's mean latency. The separate paired interface-cure gate did not pass. A matched ablation fixed the source cases, 52-candidate pool, 24-target count, model, recipe, and seed while changing target selection. On 160 new cases, base, schema-selected, model-self-selected, and VAL-selected execution reached 1, 55, 69, and 102 accepted plans. VAL exceeded self-selection by paired net +33 (p=0.0000019647), with gains in both difficulty strata. Independent semantic selection is therefore load-bearing relative to matched alternatives within this band. A complementary Phi stronger-teacher arm raised base Phi-4 from 2 to 51 accepted plans and from 35 to 80 schema-valid outputs. Earlier synthetic experiments establish teachability, cumulative learning, construction robustness, and stopping boundaries. The results support verifier-selected supervision for bounded, machine-checkable capabilities, not arbitrary planning, enterprise validity, or unrestricted self-improvement.

流程修复验证监督高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。