arXiv:2506.08647cs.CLcs.AI2025-06被引 3

用大模型摘要精炼微生物组文本,提升关系抽取效果。

Summarization for Generative Relation Extraction in the Microbiome Domain

  • 先用大模型摘要压缩文本,再生成关系,减少噪声
  • 摘要后生成式抽取在自建数据集上表现更好
  • 适合资源少的生物医学领域研究者参考

我们探索了一种针对肠道微生物组复杂且低资源生物医学领域的生成式关系抽取(RE)流程。该方法利用大语言模型(LLMs)进行文本摘要,以提炼上下文信息,随后通过指令微调的生成模型进行关系抽取。在自建语料库上的初步结果显示,摘要能有效降低噪声并引导模型,从而提升生成式关系抽取性能。然而,基于BERT的关系抽取方法仍优于生成式模型。这项持续工作展示了生成式方法在低资源专业领域中的应用潜力。

原文摘要 · Abstract (English)

We explore a generative relation extraction (RE) pipeline tailored to the study of interactions in the intestinal microbiome, a complex and low-resource biomedical domain. Our method leverages summarization with large language models (LLMs) to refine context before extracting relations via instruction-tuned generation. Preliminary results on a dedicated corpus show that summarization improves generative RE performance by reducing noise and guiding the model. However, BERT-based RE approaches still outperform generative models. This ongoing work demonstrates the potential of generative methods to support the study of specialized domains in low-resources setting.

关系抽取微生物组生成式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。