arXiv:2507.18858cs.LG2025-07被引 4

让强模型从弱模型的失败经验中学习,提升复杂决策能力。

Weak-to-Strong Generalization with Failure Trajectories: A Tree-based Approach to Elicit Optimal Policy in Strong Models

  • 用弱模型生成的失败轨迹构建树状结构,指导强模型学习。
  • 在多个任务上显著提升强模型的推理与决策表现。
  • 适合需要高效学习复杂策略的研究者和开发者。

弱到强泛化(W2SG)是一种通过弱模型监督来激发强模型全部能力的新范式。现有研究多集中于二分类等简单任务,本文将其拓展至复杂的交互式决策环境。我们通过微调强模型,利用弱模型生成的中间动作轨迹进行训练。受人类学习过程启发,提出不仅传递成功经验,也传递失败经历,使强模型能从弱模型积累的失败轨迹中学习。为进一步有效挖掘强代理潜力,我们构建了“轨迹树”——一种层次化表示结构,结合蒙特卡洛树搜索(MCTS)优化强模型。理论分析提供了方法有效性的形式化保证。实证结果表明,在多个任务领域中,该框架显著提升了推理与决策能力,验证了其可扩展性与鲁棒性。

原文摘要 · Abstract (English)

Weak-to-Strong generalization (W2SG) is a new trend to elicit the full capabilities of a strong model with supervision from a weak model. While existing W2SG studies focus on simple tasks like binary classification, we extend this paradigm to complex interactive decision-making environments. Specifically, we fine-tune a strong model with trajectories of intermediate actions generated by a weak model. Motivated by the human learning process, we propose to generalize not only success knowledge but also failure experience so that the strong model can learn from failed trajectories accumulated by weak models. To effectively and efficiently elicit the potential of strong agents, we further construct ``trajectory trees," a hierarchical representation that organizes weak model-generated action trajectories, coupled with Monte Carlo Tree Search (MCTS) to optimize the strong model. Through theoretical analysis, we provide formal guarantees for the effectiveness of our method in improving W2SG performance. Our empirical evaluations demonstrate substantial improvements in reasoning and decision-making capabilities across diverse task domains, validating the scalability and robustness of our proposed framework.

强化学习轨迹树决策优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。