构建首个大规模饮食-微生物组关联语料库,助力精准营养研究。
DiMB-RE: Mining the Scientific Literature for Diet-Microbiome Associations
- 构建含1.4万实体、4206条关系的饮食-微生物组语料库
- 微调模型在实体识别上表现良好,关系抽取准确率约44.5%
- 适合从事生物医学文本挖掘与个性化营养研究者使用
目标:从生物医学文献中构建一个标注了饮食-微生物组关联的语料库,并训练自然语言处理(NLP)模型以识别这些关联,从而加深对它们在健康与疾病中作用的理解,支持个性化营养策略。方法:我们构建了DiMB-RE,一个包含15种实体类型(如营养素、微生物)和13种关系类型(如增加、改善)的综合性语料库,涵盖165篇文献(包括30篇全文结果部分)。我们对最先进的NLP模型进行了微调与评估,用于命名实体、触发词和关系抽取,以及事实性检测。此外,我们在数据集子集上零样本和单样本设置下对两个生成式大模型(GPT-4o-mini 和 GPT-4o)进行了基准测试。结果:DiMB-RE包含14,450个实体和4,206条关系。微调后的模型在命名实体识别上表现合理(F1得分0.800),而端到端关系抽取表现中等(F1得分0.445)。使用结果部分标注提升了关系抽取效果。触发词检测影响参差不齐。生成式模型的表现低于微调模型。讨论:据我们所知,DiMB-RE是目前最大且最多样化的饮食-微生物组交互语料库。在该领域,基于DiMB-RE微调的NLP模型性能低于类似语料库,凸显了信息抽取的复杂性。误分类实体、遗漏触发词和跨句关系是关系抽取错误的主要来源。结论:DiMB-RE可作为生物医学文献挖掘的基准语料库。DiMB-RE及NLP模型已在 https://github.com/ScienceNLP-Lab/DiMB-RE 公开。
原文摘要 · Abstract (English)
Objective: To develop a corpus annotated for diet-microbiome associations from the biomedical literature and train natural language processing (NLP) models to identify these associations, thereby improving the understanding of their role in health and disease, and supporting personalized nutrition strategies. Materials and Methods: We constructed DiMB-RE, a comprehensive corpus annotated with 15 entity types (e.g., Nutrient, Microorganism) and 13 relation types (e.g., INCREASES, IMPROVES) capturing diet-microbiome associations. We fine-tuned and evaluated state-of-the-art NLP models for named entity, trigger, and relation extraction as well as factuality detection using DiMB-RE. In addition, we benchmarked two generative large language models (GPT-4o-mini and GPT-4o) on a subset of the dataset in zero- and one-shot settings. Results: DiMB-RE consists of 14,450 entities and 4,206 relationships from 165 publications (including 30 full-text Results sections). Fine-tuned NLP models performed reasonably well for named entity recognition (0.800 F1 score), while end-to-end relation extraction performance was modest (0.445 F1). The use of Results section annotations improved relation extraction. The impact of trigger detection was mixed. Generative models showed lower accuracy compared to fine-tuned models. Discussion: To our knowledge, DiMB-RE is the largest and most diverse corpus focusing on diet-microbiome interactions. NLP models fine-tuned on DiMB-RE exhibit lower performance compared to similar corpora, highlighting the complexity of information extraction in this domain. Misclassified entities, missed triggers, and cross-sentence relations are the major sources of relation extraction errors. Conclusions: DiMB-RE can serve as a benchmark corpus for biomedical literature mining. DiMB-RE and the NLP models are available at https://github.com/ScienceNLP-Lab/DiMB-RE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。