arXiv:2503.00624cs.CLcs.AI2025-03被引 11

评测DeepSeek模型在生物医学NLP中的表现,发现其命名实体识别强但事件抽取弱。

An evaluation of DeepSeek Models in Biomedical Natural Language Processing

  • 对比12个数据集,测试DeepSeek系列模型在四类生物医学任务中的表现。
  • 命名实体识别和文本分类表现接近顶尖模型,事件与关系抽取精度不足。
  • 为不同任务推荐合适模型,适合关注生物医学LLM落地的研究者参考。

大型语言模型(LLMs)的进展显著推动了生物医学自然语言处理(NLP)的发展,提升了命名实体识别、关系抽取、事件抽取和文本分类等任务的性能。在此背景下,DeepSeek系列模型在通用NLP任务中展现出潜力,但在生物医学领域的能力仍缺乏系统评估。本研究在12个数据集上,对多个DeepSeek模型(Distilled-DeepSeek-R1系列和Deepseek-LLMs)进行了评估,涵盖命名实体识别、关系抽取、事件抽取和文本分类四项关键任务,并与Llama3-8B、Qwen2.5-7B、Mistral-7B、Phi-4-14B、Gemma-2-9B等先进模型进行对比。结果表明,尽管在命名实体识别和文本分类任务中表现具有竞争力,但在事件抽取和关系抽取方面仍存在精确率-召回率权衡问题。研究提供了任务导向的模型推荐,并指明未来研究方向。该评估揭示了DeepSeek模型在生物医学NLP中的优势与局限,为后续部署与优化提供依据。

原文摘要 · Abstract (English)

The advancement of Large Language Models (LLMs) has significantly impacted biomedical Natural Language Processing (NLP), enhancing tasks such as named entity recognition, relation extraction, event extraction, and text classification. In this context, the DeepSeek series of models have shown promising potential in general NLP tasks, yet their capabilities in the biomedical domain remain underexplored. This study evaluates multiple DeepSeek models (Distilled-DeepSeek-R1 series and Deepseek-LLMs) across four key biomedical NLP tasks using 12 datasets, benchmarking them against state-of-the-art alternatives (Llama3-8B, Qwen2.5-7B, Mistral-7B, Phi-4-14B, Gemma-2-9B). Our results reveal that while DeepSeek models perform competitively in named entity recognition and text classification, challenges persist in event and relation extraction due to precision-recall trade-offs. We provide task-specific model recommendations and highlight future research directions. This evaluation underscores the strengths and limitations of DeepSeek models in biomedical NLP, guiding their future deployment and optimization.

生物医学NLP大模型评测命名实体识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。