梳理医疗大模型的语料、定制与评估,揭示数据偏见与标准缺失问题
Exploring Large Language Models in Healthcare: Insights into Corpora Sources, Customization Strategies, and Evaluation Metrics
- 分析61篇2021-2024年研究,归纳四类医疗语料来源与多方法融合策略
- 发现44项研究结合多种技术,但评估仍缺乏统一框架,结果依赖专家判断
- 强调需构建分级可信语料库与透明化模型,推动真实场景验证
本研究系统回顾了2021至2024年间61篇关于大型语言模型(LLMs)在医疗领域应用的研究,聚焦其训练语料、定制策略与评估指标。研究识别出四类语料:临床资源、文献资料、开源数据集及网络爬取数据。常见构建方法包括预训练、提示工程和检索增强生成,其中44项研究采用多种方法结合。评估指标分为流程、可用性和结果三类,结果指标又细分为模型自评与专家评估两类。研究发现语料公平性不足,导致地理、文化与社会经济因素引发偏见;对未经验证或非结构化数据的依赖凸显了整合循证临床指南的必要性。未来应发展分级语料架构,引入经审核来源与动态权重机制,并确保模型可解释性。此外,领域专用模型缺乏标准化评估框架,亟需在真实医疗环境中开展全面验证。
原文摘要 · Abstract (English)
This study reviewed the use of Large Language Models (LLMs) in healthcare, focusing on their training corpora, customization techniques, and evaluation metrics. A systematic search of studies from 2021 to 2024 identified 61 articles. Four types of corpora were used: clinical resources, literature, open-source datasets, and web-crawled data. Common construction techniques included pre-training, prompt engineering, and retrieval-augmented generation, with 44 studies combining multiple methods. Evaluation metrics were categorized into process, usability, and outcome metrics, with outcome metrics divided into model-based and expert-assessed outcomes. The study identified critical gaps in corpus fairness, which contributed to biases from geographic, cultural, and socio-economic factors. The reliance on unverified or unstructured data highlighted the need for better integration of evidence-based clinical guidelines. Future research should focus on developing a tiered corpus architecture with vetted sources and dynamic weighting, while ensuring model transparency. Additionally, the lack of standardized evaluation frameworks for domain-specific models called for comprehensive validation of LLMs in real-world healthcare settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。