arXiv:2607.03935cs.AI2026-07被引 4

让模型同时优化任务解法和辅助工具,实现自我进化。

Harness-Aware Self-Evolving: Co-Evolving Model Weights, Harness, and Task Solutions

  • 模型在多轮交互中既生成解法又修改辅助工具
  • 性能接近120B大模型,且在算法发现上达顶尖水平
  • 适合需要自动优化流程的AI系统开发者

自进化框架通常在固定辅助工具(harness)环境下优化任务解法。本文提出Harness-Aware Self-Evolving(HASE),一种基于智能体强化学习的统一框架,使单个Qwen3-8B模型能在多轮动作空间中生成任务解法或编辑选定的harness组件。HASE使该模型在文本分类任务上达到与使用Claude Code作为harness提议者的GPT-OSS-120B模型相当的性能。在alpha因子挖掘任务中,优于报告的GPT-OSS-120B基线。此外,HASE能修复不完善的评估组件,并在圆堆叠算法发现任务中收敛至当前最优表现。结果表明,HASE通过统一的智能体过程同步改进harness与解决方案。

原文摘要 · Abstract (English)

Self-evolving frameworks usually optimize task solutions while treating the surrounding harness as fixed. We introduce Harness-Aware Self-Evolving (HASE), an agentic reinforcement-learning framework in which a single model can generate task solutions or edit selected harness components in a multi-turn action space. HASE enables a single Qwen3-8B model to match the text-classification performance of a GPT-OSS-120B model that uses Claude Code as the harness proposer. In alpha factor mining, HASE outperforms the reported GPT-OSS-120B baseline. HASE also repairs imperfect evaluation components and converges to state-of-the-art performance in circle-packing algorithm discovery. These results show that HASE improves the harness and the solution through one unified agentic process.

自进化智能体强化学习模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。