用CRAFT数据集测试大模型在生物医学共指消解中的表现
BioCoref: Benchmarking Biomedical Coreference Resolution with LLMs
- 用提示工程融合领域术语和缩写信息提升模型性能
- LLaMA 8B/17B在实体增强提示下表现最优,F1更高
- 轻量级提示设计可有效增强大模型在生物医学任务中的应用
生物医学文本中的共指消解因领域专有术语复杂、指代形式模糊及长距离依赖而面临独特挑战。本文基于CRAFT语料库,全面评估生成式大语言模型(LLMs)在该领域的共指消解表现。通过四种提示策略,考察局部上下文、情境增强与领域线索(如缩写、实体词典)的使用效果。对比判别式模型SpanBERT,结果表明:尽管大模型具备较强的表面共指能力,尤其在加入领域提示后,其性能仍受长距离上下文和指代模糊性影响。值得注意的是,LLaMA 8B和17B在实体增强提示下展现出更优的精确率与F1分数,凸显轻量级提示工程在提升大模型在生物医学NLP任务中实用性的潜力。
原文摘要 · Abstract (English)
Coreference resolution in biomedical texts presents unique challenges due to complex domain-specific terminology, high ambiguity in mention forms, and long-distance dependencies between coreferring expressions. In this work, we present a comprehensive evaluation of generative large language models (LLMs) for coreference resolution in the biomedical domain. Using the CRAFT corpus as our benchmark, we assess the LLMs' performance with four prompting experiments that vary in their use of local, contextual enrichment, and domain-specific cues such as abbreviations and entity dictionaries. We benchmark these approaches against a discriminative span-based encoder, SpanBERT, to compare the efficacy of generative versus discriminative methods. Our results demonstrate that while LLMs exhibit strong surface-level coreference capabilities, especially when supplemented with domain-grounding prompts, their performance remains sensitive to long-range context and mentions ambiguity. Notably, the LLaMA 8B and 17B models show superior precision and F1 scores under entity-augmented prompting, highlighting the potential of lightweight prompt engineering for enhancing LLM utility in biomedical NLP tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。