arXiv:2502.16450cs.CL2025-02中稿 · the Symposium on I…

用可复现的流程重振文献发现,让科学假设生成更可靠。

Make Literature-Based Discovery Great Again through Reproducible Pipelines

  • 提供完整Jupyter笔记本流程,涵盖数据获取到假设评估
  • 整合多种方法并公开代码与Docker镜像,支持一键复现
  • 适合希望验证或扩展文献发现研究的科研人员

通过连接分散的科学文献,基于文献发现(LBD)方法能揭示单一领域文档无法获得的新知识和研究假设。本文聚焦于结合双关推理与LBD技术的双关式LBD方法,从可复现科学角度出发,确保LBD实验的可重复性,解决基准数据集和方法使用不一致的问题,促进协作,并推动该领域向更稳健、更具影响力的科学发现迈进。本研究的核心创新在于提供一套Jupyter Notebook,系统演示双关式LBD流程,包括数据获取、文本预处理、假设生成与评估。这些笔记实现了一系列传统LBD方法,以及作者提出的集成式、异常检测式和链接预测式方法。读者可通过开放数据集、代码复用及预配置Docker环境,获得直接上手实践的便利,保障所选LBD方法的可复现性。

原文摘要 · Abstract (English)

By connecting disparate sources of scientific literature, literature\-/based discovery (LBD) methods help to uncover new knowledge and generate new research hypotheses that cannot be found from domain-specific documents alone. Our work focuses on bisociative LBD methods that combine bisociative reasoning with LBD techniques. The paper presents LBD through the lens of reproducible science to ensure the reproducibility of LBD experiments, overcome the inconsistent use of benchmark datasets and methods, trigger collaboration, and advance the LBD field toward more robust and impactful scientific discoveries. The main novelty of this study is a collection of Jupyter Notebooks that illustrate the steps of the bisociative LBD process, including data acquisition, text preprocessing, hypothesis formulation, and evaluation. The contributed notebooks implement a selection of traditional LBD approaches, as well as our own ensemble-based, outlier-based, and link prediction-based approaches. The reader can benefit from hands-on experience with LBD through open access to benchmark datasets, code reuse, and a ready-to-run Docker recipe that ensures reproducibility of the selected LBD methods.

文献发现可复现双关推理Jupyter

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。