小模型通过微调可高效生成生物医学本体关系,突破推理瓶颈。
Benchmarking Resource-Efficient LLMs for Research Topic Ontology Generation in the Biomedical Field

- 用微调提升小模型对生物医学概念关系的识别能力。
- 微调后平均F1得分提升34.1个百分点,达68.7%。
- 适合需要低成本自动化构建本体的研究者使用。
知识组织系统如本体和分类体系是结构化科学知识的基础,但其人工构建始终是知识管理中的瓶颈。尽管大语言模型(LLMs)为自动化本体生成提供了可扩展路径,但其在识别复杂领域语义关系方面的能力仍需系统评估。本文评估了五种小型开源LLM(参数量不超过90亿)在识别生物医学概念间语义关系上的表现。为此,我们构建了MeSH-Rel-4K数据集,包含从医学主题词表(MeSH)中提取的4000个语义关系。分析了三种适配策略:标准提示、思维链提示和微调。尽管参数受限模型传统上难以处理上下文逻辑,但实验显示,针对性微调使平均F1分数提升34.1个百分点,达到68.7%。结果表明,直接微调能有效克服小模型的推理局限,为专业生物医学本体的构建与演化提供准确、自动化的方案。
原文摘要 · Abstract (English)
Knowledge Organization Systems like Ontologies and taxonomies are fundamental for structuring scientific knowledge, yet their manual curation presents a persistent bottleneck in knowledge management. While Large Language Models (LLMs) offer a scalable mechanism for automated ontology generation, their capacity to classify complex, domain-specific semantics requires systematic evaluation. In this paper, we assess the performance of five small, open-source LLMs (up to 9 billion parameters) in identifying semantic relationships between biomedical concepts. To support this evaluation, we introduce MeSH-Rel-4K, a dataset comprising 4K semantic relationships extracted from the Medical Subject Headings (MeSH). We analyse three adaptation strategies: standard prompting, Chain-of-Thought prompting, and fine-tuning. While parameter-constrained models traditionally struggle with the nuances of in-context logic, our results reveal that targeted fine-tuning increases the average F1-score by 34.1 percentage points. These results confirm that direct fine-tuning effectively exceeds the reasoning bottlenecks of smaller LLMs, providing an accurate, automated methodology for the construction and evolution of specialised biomedical ontologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。