用文献和数据联合生成科学假说,提升准确性和实用性。
Literature Meets Data: A Synergistic Approach to Hypothesis Generation
- 结合文献与数据,用大模型生成假说。
- 比纯文献或纯数据方法分别高15.75%和3.37%。
- 人类使用假说后,在识骗和识伪上准确率提升超7%。
人工智能有望变革科学流程,包括假说生成。现有方法可分为理论驱动与数据驱动两类,二者虽均能生成新颖且合理的假说,但能否互补仍未知。为此,我们提出首个融合文献洞察与数据的LLM假说生成方法。在五个不同数据集上验证,该方法优于其他基线:比少样本方法高8.97%,比纯文献方法高15.75%,比纯数据方法高3.37%。此外,我们首次开展人工评估,检验大模型生成假说在欺骗检测与AI生成内容识别两个挑战任务中的辅助价值。结果显示,人类在两项任务上的准确率分别提升7.44%和14.19%。这些发现表明,融合文献与数据的方法可构建更全面、细致的假说生成框架,为科学研究开辟新路径。
原文摘要 · Abstract (English)
AI holds promise for transforming scientific processes, including hypothesis generation. Prior work on hypothesis generation can be broadly categorized into theory-driven and data-driven approaches. While both have proven effective in generating novel and plausible hypotheses, it remains an open question whether they can complement each other. To address this, we develop the first method that combines literature-based insights with data to perform LLM-powered hypothesis generation. We apply our method on five different datasets and demonstrate that integrating literature and data outperforms other baselines (8.97\% over few-shot, 15.75\% over literature-based alone, and 3.37\% over data-driven alone). Additionally, we conduct the first human evaluation to assess the utility of LLM-generated hypotheses in assisting human decision-making on two challenging tasks: deception detection and AI generated content detection. Our results show that human accuracy improves significantly by 7.44\% and 14.19\% on these tasks, respectively. These findings suggest that integrating literature-based and data-driven approaches provides a comprehensive and nuanced framework for hypothesis generation and could open new avenues for scientific inquiry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。