用猜规则游戏测试大模型的科学推理能力,发现能主动证伪的模型表现更好。
FALSIFYBENCH: Evaluating Inductive Reasoning in LLMs with Rule Discovery Games

- 设计类沃森2-4-6任务,让模型通过试错发现隐藏规则
- 仅少数模型接近最优,主动寻找反例者显著领先
- 揭示了模型在假设空间中探索失败的可识别模式
大型语言模型(LLMs)越来越多地被部署为科学任务中的自主代理。然而,这些系统是否具备科学发现所需的归纳推理能力仍不明确。本文提出FALSIFYBENCH,一个受经典沃森2-4-6任务启发的评估框架,要求代理通过迭代提出示例并接收反馈来发现隐藏语义属性。该任务捕捉了科学推理的核心要素:假设生成、证据收集及对确认与否定证据的信念修正。我们在12个不同模型家族和规模的LLM上进行评估,发现推理型模型普遍优于指令微调模型,但无一接近最优性能。成功的关键在于负向测试能力:主动寻求证伪假设的模型始终优于仅寻求验证的模型。此外,细致的逐轮分析揭示,失败与模型在假设空间中导航的特定模式密切相关。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed as autonomous agents in scientific tasks. Yet whether these systems can effectively engage in forms of inductive reasoning relevant to scientific discovery remains an open question. In this work, we introduce FALSIFYBENCH, an evaluation framework for hypothesis-driven reasoning inspired by the classic Wason 2-4-6 task, in which agents must discover hidden semantic properties by iteratively proposing examples and receiving feedback. This task captures key elements of scientific reasoning: hypothesis generation, evidence gathering, and belief revision in response to both confirming and disconfirming evidence. Our evaluation of 12 LLMs across model families and scales shows that reasoning models are generally stronger scientific reasoners than instruction-tuned models, although no model comes close to optimal performance. The primary driver of success is the capacity for negative testing: models that actively seek to falsify their hypotheses consistently outperform those that primarily seek confirmation. Moreover, a fine-grained turn-level analysis, neglected in previous work, reveals that failure is tied to identifiable patterns in how models navigate the hypothesis space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。