用结构化数据训练模型生成更新颖可行的科学假说。
Sparks of Science: Hypothesis Generation Using Structured Paper Data
- 构建包含5500对的假说生成数据集,采用'比特-火花-反转'框架。
- 微调后模型生成假说的新颖性、可行性与整体质量显著提升。
- 适合对科学创新、AI辅助研究感兴趣的开发者和研究人员。
生成新颖且具有创意的科学假说是实现通用人工智能的核心。大语言与推理模型有望助力科学假说的系统性创建、筛选与验证。然而,当前基础模型常难以生成既新颖又可行的科学构想。原因之一是缺乏将科学假说生成(SHG)定义为自然语言生成(NLG)任务的专用数据集。本文提出HypoGen,首个约5500个结构化问题-假说对的数据集,源自顶级计算机科学会议,采用“比特-火花-反转”(Bit-Flip-Spark)架构:比特为传统假设,火花为关键洞察或概念突破,反转为由此产生的反向提案。HypoGen独特地整合了显式的思维链(Chain-of-Reasoning)组件,反映从比特到反转的认知过程。我们证明,将假说生成建模为条件语言建模任务,使用比特-火花-反转及思维链进行微调,在推理时仅提供比特,可显著提升生成假说的整体质量。评估采用自动化指标与大模型裁判排名。实验表明,基于HypoGen微调后,生成假说在新颖性、可行性与整体质量上均有提升。HypoGen数据集已公开于 huggingface.co/datasets/UniverseTBD/hypogen-dr1。
原文摘要 · Abstract (English)
Generating novel and creative scientific hypotheses is a cornerstone in achieving Artificial General Intelligence. Large language and reasoning models have the potential to aid in the systematic creation, selection, and validation of scientifically informed hypotheses. However, current foundation models often struggle to produce scientific ideas that are both novel and feasible. One reason is the lack of a dedicated dataset that frames Scientific Hypothesis Generation (SHG) as a Natural Language Generation (NLG) task. In this paper, we introduce HypoGen, the first dataset of approximately 5500 structured problem-hypothesis pairs extracted from top-tier computer science conferences structured with a Bit-Flip-Spark schema, where the Bit is the conventional assumption, the Spark is the key insight or conceptual leap, and the Flip is the resulting counterproposal. HypoGen uniquely integrates an explicit Chain-of-Reasoning component that reflects the intellectual process from Bit to Flip. We demonstrate that framing hypothesis generation as conditional language modelling, with the model fine-tuned on Bit-Flip-Spark and the Chain-of-Reasoning (and where, at inference, we only provide the Bit), leads to improvements in the overall quality of the hypotheses. Our evaluation employs automated metrics and LLM judge rankings for overall quality assessment. We show that by fine-tuning on our HypoGen dataset we improve the novelty, feasibility, and overall quality of the generated hypotheses. The HypoGen dataset is publicly available at huggingface.co/datasets/UniverseTBD/hypogen-dr1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。