arXiv:2412.18154q-bio.GNcs.AI2024-12中稿 · BIBM 2024被引 2

用大模型自动提取基因研究摘要,提升文献信息整合效率。

GeneSUM: Large Language Model-based Gene Summary Extraction

  • 两阶段流程:先筛选文献去重,再微调大模型生成摘要。
  • 大模型显著提升基因信息整合能力,支持高效科研决策。
  • 适合生物医学研究者快速掌握基因功能最新进展。

生物医学研究新领域持续扩展,带来大量关于基因及其功能的信息。知识的快速增长为科学发现提供了前所未有的机遇,同时也给研究人员及时跟进最新进展带来了巨大挑战。其中一大难题是处理海量文献以提取关键基因信息,这一过程耗时且繁琐。为此,我们提出 GeneSUM,一种基于大语言模型(LLM)的两阶段自动化基因摘要提取框架。该方法首先检索并去除目标基因相关文献中的冗余内容,随后微调 LLM 以优化和精简摘要生成过程。通过大量实验验证,结果表明,大语言模型显著提升了基因特异性信息的整合能力,从而在持续研究中实现更高效的决策支持。

原文摘要 · Abstract (English)

Emerging topics in biomedical research are continuously expanding, providing a wealth of information about genes and their function. This rapid proliferation of knowledge presents unprecedented opportunities for scientific discovery and formidable challenges for researchers striving to keep abreast of the latest advancements. One significant challenge is navigating the vast corpus of literature to extract vital gene-related information, a time-consuming and cumbersome task. To enhance the efficiency of this process, it is crucial to address several key challenges: (1) the overwhelming volume of literature, (2) the complexity of gene functions, and (3) the automated integration and generation. In response, we propose GeneSUM, a two-stage automated gene summary extractor utilizing a large language model (LLM). Our approach retrieves and eliminates redundancy of target gene literature and then fine-tunes the LLM to refine and streamline the summarization process. We conducted extensive experiments to validate the efficacy of our proposed framework. The results demonstrate that LLM significantly enhances the integration of gene-specific information, allowing more efficient decision-making in ongoing research.

基因摘要大模型文献挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。