arXiv:2511.19739cs.CLcs.LG2025-11

对比10种心病文本嵌入模型,发现小模型更优且省资源

Comparative Analysis of LoRA-Adapted Embedding Models for Clinical Cardiology Text Representation

  • 用低秩适配微调10种模型,基于10万+心脏病文本对
  • BioLinkBERT表现最佳,分离度达0.510,远超大模型
  • 适合临床NLP开发,代码数据全公开可复现

领域特定文本嵌入对临床自然语言处理至关重要,但不同模型架构的系统性比较仍有限。本研究在来自权威医学教材的106,535条心脏病文本对上,通过低秩适配(LoRA)微调评估了十种基于Transformer的嵌入模型。结果表明,编码器型架构,尤其是BioLinkBERT,在领域特定性能上表现最优(分离度得分0.510),同时所需计算资源显著少于更大规模的解码器型模型。研究挑战了大语言模型必然产生更好领域嵌入的假设,为临床NLP系统开发提供了实用指导。所有模型、训练代码及评估数据集均公开,支持医学信息学领域的可复现研究。

原文摘要 · Abstract (English)

Domain-specific text embeddings are critical for clinical natural language processing, yet systematic comparisons across model architectures remain limited. This study evaluates ten transformer-based embedding models adapted for cardiology through Low-Rank Adaptation (LoRA) fine-tuning on 106,535 cardiology text pairs derived from authoritative medical textbooks. Results demonstrate that encoder-only architectures, particularly BioLinkBERT, achieve superior domain-specific performance (separation score: 0.510) compared to larger decoder-based models, while requiring significantly fewer computational resources. The findings challenge the assumption that larger language models necessarily produce better domain-specific embeddings and provide practical guidance for clinical NLP system development. All models, training code, and evaluation datasets are publicly available to support reproducible research in medical informatics.

心病文本LoRA嵌入模型临床NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。