用多智能体辩论机制生成更真实可信的科学创意。
Debate as Reward: A Multi-Agent Reward System for Scientific Ideation via RL Post-Training

- 设计多智能体辩论作为奖励函数,避免虚假创新。
- 在ICLR-320数据集上显著优于现有方法。
- 适合需要高质量科学创意的科研人员使用。
大型语言模型在自动化科学创意生成方面展现出潜力,但当前依赖迭代提示或复杂多智能体架构的方法常因幻觉或计算效率低下而受限。将强化学习应用于这一开放领域的一个关键瓶颈是奖励滥用——模型利用不完善的评估代理来最大化得分,而非产生真正的科学创新。为此,我们提出首个专为高质量科学创意生成设计的强化学习框架。该框架采用多智能体奖励函数作为评判者,将方法验证与实现细节解耦,提供对奖励滥用具有鲁棒性的严格二元奖励。为有效优化这一稀疏信号,我们采用无偏的组相对策略优化变体,以缓解人为长度偏差。训练基于ICLR-320数据集,该数据集从ICLR 2024会议论文中提取了320对问题-解决方案。实验表明,我们的框架在专家评估的新颖性、可行性与有效性指标上均显著优于现有最先进基线。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated potential in automating scientific ideation, yet current approaches relying on iterative prompting or complex multi-agent architectures often suffer from hallucination or computational inefficiency. A critical bottleneck in applying Reinforcement Learning (RL) to this open-ended domain is reward hacking -- where models exploit imperfect evaluation proxies to maximize scores without producing genuine scientific innovation. To address these limitations, we propose an RL framework explicitly tailored for high-quality scientific idea generation. We propose the first multi-agent reward function designed to serve as a judge, decoupling methodological validation from implementation details while providing strict binary rewards that are robust to reward hacking. To effectively optimize against this sparse signal, we utilize an unbiased variant of Group Relative Policy Optimization to mitigate artificial length bias. We grounded our training in ICLR-320, a curated dataset of problem-solution pairs extracted from ICLR 2024 proceedings. Experiments demonstrate that our framework significantly outperforms state-of-the-art baselines across expert-evaluated metrics of novelty, feasibility, and effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。