构建系统化基准测试,评估大模型生成科学假设的能力。
HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation
- 设计多维度基准HypoBench,覆盖真实与合成任务
- 合成数据下最佳模型仅发现38.8%真实假设,难度上升性能骤降
- 为科学发现类AI系统提供可复现的评估工具
大型语言模型在假设生成领域受到广泛关注,但核心问题仍存:何为优质假设?如何系统评估生成方法?为此,我们提出HypoBench,一个新型基准,用于评估大模型及假设生成方法在实际效用、泛化能力与假设发现率等方面的性能。该基准包含7个真实世界任务和5个合成任务,涵盖194个不同数据集。我们评估了四种顶尖大模型与六种现有假设生成方法的组合表现。结果表明,现有方法能在数据中发现有效且新颖的模式。然而,在合成数据上,随着任务难度增加,性能显著下降,最优模型与方法仅能恢复38.8%的真实假设,说明仍有较大改进空间。这些发现揭示了假设生成中的挑战,并证明HypoBench是提升辅助科学发现的人工智能系统的重要资源。
原文摘要 · Abstract (English)
There is growing interest in hypothesis generation with large language models (LLMs). However, fundamental questions remain: what makes a good hypothesis, and how can we systematically evaluate methods for hypothesis generation? To address this, we introduce HypoBench, a novel benchmark designed to evaluate LLMs and hypothesis generation methods across multiple aspects, including practical utility, generalizability, and hypothesis discovery rate. HypoBench includes 7 real-world tasks and 5 synthetic tasks with 194 distinct datasets. We evaluate four state-of-the-art LLMs combined with six existing hypothesis-generation methods. Overall, our results suggest that existing methods are capable of discovering valid and novel patterns in the data. However, the results from synthetic datasets indicate that there is still significant room for improvement, as current hypothesis generation methods do not fully uncover all relevant or meaningful patterns. Specifically, in synthetic settings, as task difficulty increases, performance significantly drops, with best models and methods only recovering 38.8% of the ground-truth hypotheses. These findings highlight challenges in hypothesis generation and demonstrate that HypoBench serves as a valuable resource for improving AI systems designed to assist scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。