arXiv:2412.14626cs.CLcs.AI2024-12被引 2

用动态控制生成更符合科研标准的研究创意。

LDC: Learning to Generate Research Idea with Dynamic Control

  • 分两阶段训练:先微调学模式,再强化学习优化多维质量。
  • 在新颖性、可行性、有效性间动态平衡,生成高质量创意。
  • 适合需要持续产出优质研究想法的学者或团队使用。

大型语言模型在自动化科研创意生成方面展现出潜力。现有方法多依赖提示工程,常生成与专家标准(新颖性、可行性、有效性)不符的创意,且三者存在固有权衡。为此,我们提出首个结合监督微调(SFT)和可控强化学习(RL)的两阶段框架。在SFT阶段,模型从论文与其后续创意对中学习基础模式;在RL阶段,多维度奖励模型基于细粒度反馈评估并优化模型表现。推理时,由句级解码器协调的维度控制器实现上下文感知的动态引导。实验表明,该框架能有效平衡三者权衡,生成高质量研究创意。

原文摘要 · Abstract (English)

Recent advancements in large language models (LLMs) have demonstrated their potential in automating the scientific research ideation. Existing approaches primarily focus on prompting techniques, often producing ideas misaligned with expert standards - novelty, feasibility, and effectiveness, which are widely recognized by the research community as the three key subdimensions of high-quality ideas. Also, balancing these dimensions remains challenging due to their inherent trade-offs. To address these limitations, we propose the first framework that employs a two-stage approach combining Supervised Fine-Tuning (SFT) and controllable Reinforcement Learning (RL) for the task. In the SFT stage, the model learns foundational patterns from pairs of research papers and their corresponding follow-up ideas. In the RL stage, multi-dimensional reward models guided by fine-grained feedback evaluate and optimize the model across key dimensions. During inference, dimensional controllers coordinated by a sentence-level decoder enable dynamic context-aware steering of the idea generation process. Our framework provides a balanced approach to research idea generation, achieving high-quality outcomes in the experiment by dynamically navigating the trade-offs among novelty, feasibility, and effectiveness.

科研创新大模型生成控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。