arXiv:2505.15024cs.CL2025-05被引 2

分析大模型如何从网络数据学临床知识,发现数据与实际使用存在严重不匹配。

Diagnosing our datasets: How does my language model learn clinical information?

  • 构建新数据集MedLingo,评估模型对医学术语的理解能力。
  • 医学术语在训练数据中出现频率与模型表现正相关,但真实临床用语常缺失。
  • 大量未经证实的医疗说法出现在训练数据中,可能被模型错误复述。

大型语言模型(LLMs)在多种临床自然语言处理任务中表现良好,尽管未直接在电子健康记录(EHR)数据上训练。本文通过两个关键但研究不足的角度考察主流开源LLMs如何从大规模挖掘语料中学习临床信息:(1) 对医学术语的解读能力,这是理解真实临床笔记的基础;(2) 对无依据医疗主张的响应行为。我们分析了相关临床信息在预训练语料中的出现频率、预训练数据构成与模型输出的关系,以及这些数据的来源。为隔离术语理解能力,我们在新数据集MedLingo上评估了LLMs。结果表明,主要预训练语料中医学术语提及频率与模型性能呈正相关。然而,许多在临床笔记中频繁出现的术语在预训练语料中极少出现,暴露出数据可用性与真实使用间的严重错配。类似地,相当一部分文档支持有争议的医疗主张,这些主张随后被模型复述。最后,我们对在线来源中医学术语和未经证实医疗主张的类型进行了分类与分析,对未来的数据集构建具有重要启示。

原文摘要 · Abstract (English)

Large language models (LLMs) have performed well across various clinical natural language processing tasks, despite not being directly trained on electronic health record (EHR) data. In this work, we examine how popular open-source LLMs learn clinical information from large mined corpora through two crucial but understudied lenses: (1) their interpretation of clinical jargon, a foundational ability for understanding real-world clinical notes, and (2) their responses to unsupported medical claims. For both use cases, we investigate the frequency of relevant clinical information in their corresponding pretraining corpora, the relationship between pretraining data composition and model outputs, and the sources underlying this data. To isolate clinical jargon understanding, we evaluate LLMs on a new dataset MedLingo. Unsurprisingly, we find that the frequency of clinical jargon mentions across major pretraining corpora correlates with model performance. However, jargon frequently appearing in clinical notes often rarely appears in pretraining corpora, revealing a mismatch between available data and real-world usage. Similarly, we find that a non-negligible portion of documents support disputed claims that can then be parroted by models. Finally, we classified and analyzed the types of online sources in which clinical jargon and unsupported medical claims appear, with implications for future dataset composition.

临床NLP大模型数据偏差医学术语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。