测试大模型在不确定科学问题中生成多样假设的能力。
HypoSpace: A Diagnostic Benchmark for Set-Valued Hypothesis Generation under Underdetermination and Sublinear Coverage Bounds
- 将LLM视为有限假设空间的采样器,评估其有效性、独特性和覆盖度。
- 模型在复杂场景下虽保持高有效性,但独特性和覆盖度显著下降。
- 适合研究科学推理、模型可解释性与多假设生成的学者使用。
许多科学问题存在不确定性:多个不同假设均能与相同观测结果一致。在此类情境下,有效推断不仅需生成合理解释,还需系统探索并覆盖所有可接受的假设空间。我们提出HypoSpace,一个将大语言模型(LLMs)视作有限假设空间采样器的评测基准,从有效性(Validity)、唯一性(Uniqueness)和恢复率(Recovery)三方面进行评估。HypoSpace涵盖三个结构化领域(因果图推断、引力约束3D体素重建、布尔基因互作建模),每个领域均有确定性验证器和精确可枚举解空间,并包含真实世界锚定案例研究。实证结果显示,随着可接受假设空间增大或组合复杂度提升,模型表现出能力与规模相关的覆盖失败:虽维持高有效性,但唯一性和恢复率显著降低。进一步分析表明,分层解码策略可部分缓解该现象,证明HypoSpace作为集合值推断诊断基准的有效性。代码已开源:https://github.com/CTT-Pavilion/_HypoSpace。
原文摘要 · Abstract (English)
Many scientific problems are underdetermined: multiple distinct hypotheses are equally consistent with the same observations. In such settings, effective inference requires not only producing valid explanations, but also systematically exploring and covering the admissible hypothesis set. We introduce HypoSpace, a benchmark that treats large language models (LLMs) as samplers over finite hypothesis spaces and evaluates them on three metrics: Validity, Uniqueness, and Recovery. HypoSpace spans three structured domains (causal graph inference, gravity-constrained 3D voxel reconstruction, and Boolean genetic interaction modeling) with deterministic validators and exactly enumerable solution spaces, plus real-world anchored case studies. Empirically, HypoSpace reveals a capability- and scale-dependent coverage failure: models can maintain high Validity while exhibiting reduced Uniqueness and Recovery as admissible hypothesis spaces become larger or more combinatorial. We further show that the analysis on stratified decoding partially mitigates this collapse, demonstrating HypoSpace's utility as a diagnostic benchmark for set-valued inference. Code is available at: https://github.com/CTT-Pavilion/_HypoSpace.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。