用进化强化学习让AI自主探索科学难题,解题能力超越GPT-4o。
Helix: Evolutionary Reinforcement Learning for Open-Ended Scientific Problem Solving
- 构建多层次进化框架,结合上下文经验扩展解空间。
- 在圆堆叠任务中达成半径和2.63598308的最优结果。
- 适合需要持续创新的开放性科学问题求解场景。
具备推理能力的大语言模型在解决复杂科学问题上展现出日益增长的潜力。然而,这类任务具有高度领域性、无边界且开放,需在广阔灵活的解空间中进行探索。现有方法,无论是纯学习驱动还是依赖精心设计的工作流,常面临探索效率低和泛化能力差的问题。为此,我们提出HELIX——一种基于上下文经验的分层进化强化学习框架。该框架引入两项关键创新:(i) 通过上下文学习构建多样化且高质量的候选解池,扩大探索范围;(ii) 采用强化学习进行迭代策略优化,逐步提升解的质量。两者协同实现更优解的发现。在圆堆叠任务中,仅使用140亿参数模型即达到半径总和2.63598308的当前最优结果。在标准机器学习基准测试中,相较于经过精心设计的GPT-4o工作流,HELIX在Adult与Bank Marketing数据集上平均F1提升5.95点。
原文摘要 · Abstract (English)
Large language models (LLMs) with reasoning abilities have demonstrated growing promise for tackling complex scientific problems. Yet such tasks are inherently domain-specific, unbounded and open-ended, demanding exploration across vast and flexible solution spaces. Existing approaches, whether purely learning-based or reliant on carefully designed workflows, often suffer from limited exploration efficiency and poor generalization. To overcome these challenges, we present HELIX -- a Hierarchical Evolutionary reinforcement Learning framework with In-context eXperiences. HELIX introduces two key novelties: (i) a diverse yet high-quality pool of candidate solutions that broadens exploration through in-context learning, and (ii) reinforcement learning for iterative policy refinement that progressively elevates solution quality. This synergy enables the discovery of more advanced solutions. On the circle packing task, HELIX achieves state-of-the-art result with a sum of radii of 2.63598308 using only a 14B model. Across standard machine learning benchmarks, HELIX further surpasses GPT-4o with a carefully engineered pipeline, delivering an average F1 improvement of 5.95 points on the Adult and Bank Marketing datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。