arXiv:2607.15766cs.CL2026-07

测试大模型在无结论前提下自主提出可验证假设的能力。

Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery

论文配图:Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery
图 1 · 摘自论文原文
  • 构建新任务PHD,让模型从模糊证据中生成可检验的假设空间。
  • 创建988例跨领域的基准数据集HypoData和多维评估框架HypoEval。
  • 发现不同模型表现差异显著,评估结果与专家意见高度一致。

大型语言模型(LLMs)擅长回答预设问题,但在开放、无结论的探索阶段能力尚未被充分衡量。本文提出前瞻性假设发现(PHD)任务,要求模型从不完整证据(如异常观测和碎片化记录)中自主构建有依据、可区分、可验证的假设空间以指导后续调查。为此,我们构建了包含988个案例的HypoData基准数据集,覆盖六个科学与分析领域,并设计了HypoEval评估框架,用于评判开放式假设集合。为规模化构建数据,我们提出回溯上下文回归方法,通过“锻造-审计”流程从已完成的专家文档中移除明确结论、目标假设及事后因果归因,保留事实基础。由于PHD存在多个有效输出,HypoEval结合双向成对判断与Bradley-Terry-Davidson聚合进行排序,并采用六维评分标准进行诊断。对15个前沿大模型的实验显示,模型能力分层明显,且结构化分析技能影响各异:部分低性能模型在该基准上提升,但也有系统出现退步,包括一个顶级模型。相比绝对评分,竞技场式评估能更精细区分模型表现,其聚合排名与人类专家及独立评审高度一致。结果支持将PHD作为衡量大模型在未给出最终结论时规划研究方向的重要指标。代码与数据已开源于github.com/SKYLENAGE-AI/HypoArena。

原文摘要 · Abstract (English)

Large language models (LLMs) excel at answering pre-specified questions, yet their ability to navigate the open-ended, pre-conclusion stage of discovery remains largely unmeasured. We introduce Prospective Hypothesis Discovery (PHD), which asks models to autonomously construct grounded, discriminative, and testable hypothesis spaces from inconclusive evidence, including anomalous observations and fragmented records, to guide subsequent investigation. To evaluate this capability, we introduce HypoArena, comprising HypoData, a benchmark of 988 cases across six scientific and analytical domains, and HypoEval, an evaluation framework for open-ended hypothesis sets. To construct HypoData at scale, we propose Retrospective Context Regression, a Forge--Audit pipeline that reconstructs pre-conclusion contexts from completed expert documents by removing explicit conclusions, target hypotheses, and retrospective causal attributions while preserving the factual substrate. Because PHD admits multiple valid outputs, HypoEval combines bidirectional pairwise judgments with Bradley--Terry--Davidson aggregation for ranking and six-dimensional rubric scoring for diagnosis. Experiments on 15 frontier LLMs reveal clear capability stratification and model-dependent effects of structured analytical skills, with gains for several lower-performing models on HypoArena but regressions for other systems, including a top-performing model. Compared with absolute rubric scoring, arena evaluation resolves finer-grained differences among models, with aggregated rankings showing strong agreement with human experts and an independent judge. Together, these results support treating PHD as a distinct target for evaluating how LLMs formulate investigative directions when final conclusions are withheld. Our code and data are publicly available at github.com/SKYLENAGE-AI/HypoArena and github.com/SKYLENAGE-AI/HypoArena.

假设发现大模型评估开放探究基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。