测试大模型在逐步获取信息时生成科学假设的能力,评估其创新能力与推理可靠性。
ProjectionBench: Evaluating Scientific Hypothesis Generation in LLMs Under Progressive Information Disclosure

- 从零信息开始,逐步披露实验细节,让模型生成针对科研问题的假设。
- GPT-5.4在信息极少时仍保持0.7的F1分数与真实结论对齐,表现最优。
- 适合关注AI辅助科研、科学发现智能系统的研究者和开发者。
科学发现是高度创造性和不确定的过程,需要超越已有知识记忆的推理能力。尽管已有多个基准评估大语言模型在多跳检索任务中的表现,但其在真正科学发现中所需的创新推理能力仍缺乏有效测评。本文提出ProjectionBench框架,从原始问题出发,逐步披露技术细节,评估模型在不同信息阶段生成假设的能力。每阶段要求模型提出能回应研究问题的假设,并通过原子性陈述的语义相似度自动评估与原文结论的差异。该渐进式评估可衡量模型在信息有限时的创新性与充分信息下的严谨推理能力,两者均对构建下一代人工智能科学家系统至关重要。我们在生物活性材料、机械材料和纳米材料领域共45篇论文上评估了GPT-5、GPT-5.4、Gemini 2.5 pro和Gemini 3.1 pro preview。结果表明,GPT-5.4和Gemini 3.1 pro相比前代有提升,其中GPT-5.4在极简上下文下仍保持0.7的F1得分,与真实结论高度一致。
原文摘要 · Abstract (English)
Scientific discovery is an inherently creative and uncertain process, requiring reasoning beyond the recall of known knowledge. While many benchmarks have been proposed to evaluate large language model (LLM) performance on deep research tasks via multi-hop retrieval, their innovative reasoning abilities essential for true scientific discovery remain largely untested. We introduce a benchmark framework for evaluating model performance in scientific discovery and reasoning, building up from a raw problem to the classical null hypothesis test. In our framework, models initially receive only the topic and research question from a recent paper, with technical details progressively revealed. At each stage of information disclosure, the model is tasked with generating hypotheses that address the research question, which is compared with the conclusions from the original paper and evaluated via automated semantic similarity of constituent atomic claims. This progressive evaluation of semantic divergence from ground-truth conclusions enables assessment of a model's innovativeness (under minimal information) to grounded reasoning capabilities (under full experimental details), both critical for using LLMs for scientific discovery purposes. Our framework provides a foundation for systematically evaluating scientific reasoning and discovery capabilities in LLMs, crucial for advancing the development of next-generation AI scientist/co-scientist systems. Specifically, here we evaluate GPT-5, GPT-5.4, Gemini 2.5 pro, and Gemini 3.1 pro preview across 45 papers spanning bioactive materials, mechanical materials, and nanomaterials. We find that GPT-5.4 and Gemini 3.1 pro outperform their previous generation counterparts as expected, and GPT-5.4 in particular maintains 0.7 F1 score alignment with ground truth conclusions even under minimal context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。