测试大模型判断科学假设是否可行的能力,发现结果证据比实验描述更可靠。
Experiments or Outcomes? Probing Scientific Feasibility in Large Language Models
- 将科学可行性评估视为诊断推理任务,通过假设与证据判断可行与否。
- 有结果证据时准确率显著高于仅靠内部知识,实验描述则易因信息缺失降效。
- 适合关注科学推理可信度的研究者,尤其在模型依赖外部证据场景。
科学可行性评估旨在判断一个主张是否符合已有知识,并且是否存在实验证据支持或反驳它。本文将可行性评估建模为一项诊断推理任务:给定一个假设,模型需预测其是否可行并提供理由。我们在受控知识条件下评估大语言模型(LLMs)——仅提供假设、仅提供实验描述、仅提供结果、或同时提供实验与结果——并通过逐步移除实验和/或结果上下文来探测鲁棒性。在多个LLM和两个数据集上,结果显示:结果证据通常比实验描述更可靠。结果能提升准确率,超越模型自身知识水平;而实验文本则表现脆弱,当上下文不完整时可能降低性能。这些发现明确了实验证据在基于LLM的可行性评估中何时有益、何时引入不稳定性。
原文摘要 · Abstract (English)
Scientific feasibility assessment asks whether a claim is consistent with established knowledge and whether experimental evidence could support or refute it. We frame feasibility assessment as a diagnostic reasoning task in which, given a hypothesis, a model predicts feasible or infeasible and justifies its decision. We evaluate large language models (LLMs) under controlled knowledge conditions (hypothesis-only, with experiments, with outcomes, or both) and probe robustness by progressively removing portions of the experimental and/or outcome context. Across multiple LLMs and two datasets, providing outcome evidence is generally more reliable than providing experiment descriptions. Outcomes tend to improve accuracy beyond what internal knowledge alone provides, whereas experimental text can be brittle and may degrade performance when the context is incomplete. These findings clarify when experimental evidence benefits LLM-based feasibility assessment and when it introduces fragility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。