首个英土双语临床关系抽取评估,验证提示工程优于微调。
Relation Extraction Capabilities of LLMs on Clinical Text: A Bilingual Evaluation for English and Turkish
- 构建首个英土平行临床关系抽取数据集,支持跨语言评估。
- 提示方法整体优于微调模型,英语表现高于土耳其语。
- 提出关系感知检索法,显著提升少资源语言性能。
非英语临床信息抽取标注数据稀缺,制约了以英语为主的大型语言模型(LLM)在其他语言中的评估。本研究首次系统性地对英语与土耳其语的临床关系抽取(RE)任务开展双语评估。为此,我们基于2010年i2b2/VA关系分类语料库构建并精心整理出首个英土平行临床RE数据集。我们系统评估了多种提示策略,包括多示例学习(ICL)和思维链(CoT)方法,并与纯微调基线模型(如PURE)进行对比。此外,我们提出一种基于对比学习的关系感知检索(RAR)方法,专门捕捉句级和关系级语义。结果表明,提示类方法持续优于传统微调模型。所有评估模型中,英语表现均优于土耳其语。在ICL方法中,RAR表现最佳,Gemini 1.5 Flash在英语上达到0.906的micro-F1,在土耳其语上为0.888;当结合深度思考提示与DeepSeek-V3模型时,英语F1进一步提升至0.918。这些发现凸显高质量示范检索的重要性,也证明先进检索与提示技术能有效弥补临床自然语言处理中的资源差距。
原文摘要 · Abstract (English)
The scarcity of annotated datasets for clinical information extraction in non-English languages hinders the evaluation of large language model (LLM)-based methods developed primarily in English. In this study, we present the first comprehensive bilingual evaluation of LLMs for the clinical Relation Extraction (RE) task in both English and Turkish. To facilitate this evaluation, we introduce the first English-Turkish parallel clinical RE dataset, derived and carefully curated from the 2010 i2b2/VA relation classification corpus. We systematically assess a diverse set of prompting strategies, including multiple in-context learning (ICL) and Chain-of-Thought (CoT) approaches, and compare their performance to fine-tuned baselines such as PURE. Furthermore, we propose Relation-Aware Retrieval (RAR), a novel in-context example selection method based on contrastive learning, that is specifically designed to capture both sentence-level and relation-level semantics. Our results show that prompting-based LLM approaches consistently outperform traditional fine-tuned models. Moreover, evaluations for English performed better than their Turkish counterparts across all evaluated LLMs and prompting techniques. Among ICL methods, RAR achieves the highest performance, with Gemini 1.5 Flash reaching a micro-F1 score of 0.906 in English and 0.888 in Turkish. Performance further improves to 0.918 F1 in English when RAR is combined with a structured reasoning prompt using the DeepSeek-V3 model. These findings highlight the importance of high-quality demonstration retrieval and underscore the potential of advanced retrieval and prompting techniques to bridge resource gaps in clinical natural language processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。