arXiv:2512.04834cs.AIcs.CL2025-12被引 5

测试开源大模型在意大利医疗文本上的零样本理解能力,发现表现差异大。

Are LLMs Truly Multilingual? Exploring Zero-Shot Multilingual Capability of LLMs for Information Retrieval: An Italian Healthcare Use Case

  • 用零样本方式测试多语言大模型对意大利语病历的理解能力。
  • 部分模型在本地部署下表现不佳,跨疾病泛化能力弱。
  • 适合关注医疗AI多语言落地挑战的研究者和开发者。

大型语言模型(LLMs)已成为人工智能与自然语言处理领域的关键话题,在医疗、金融、教育和营销等领域推动客户服务优化、任务自动化、洞察提供、诊断改进及个性化学习体验。从临床记录中提取信息是数字医疗中的关键任务。尽管过去传统NLP技术曾被用于此,但受限于临床语言的复杂性、变异性及高内含语义,常难以应对自由文本。近年来,大型语言模型因其强大的人类语言理解与生成能力,成为该领域的有力工具。本文探讨开源多语言LLMs在理解意大利语电子健康记录(EHR)并实时提取信息方面的零样本能力。在针对共病提取的详细实验中,结果表明:部分模型在零样本、本地部署环境下表现不佳,且在不同疾病间的泛化性能存在显著差异,相较于原生模式匹配和人工标注结果仍有明显差距。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have become a key topic in AI and NLP, transforming sectors like healthcare, finance, education, and marketing by improving customer service, automating tasks, providing insights, improving diagnostics, and personalizing learning experiences. Information extraction from clinical records is a crucial task in digital healthcare. Although traditional NLP techniques have been used for this in the past, they often fall short due to the complexity, variability of clinical language, and high inner semantics in the free clinical text. Recently, Large Language Models (LLMs) have become a powerful tool for better understanding and generating human-like text, making them highly effective in this area. In this paper, we explore the ability of open-source multilingual LLMs to understand EHRs (Electronic Health Records) in Italian and help extract information from them in real-time. Our detailed experimental campaign on comorbidity extraction from EHR reveals that some LLMs struggle in zero-shot, on-premises settings, and others show significant variation in performance, struggling to generalize across various diseases when compared to native pattern matching and manual annotations.

大模型医疗AI多语言信息抽取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。