构建评估框架,验证大模型在医学表型识别中的效果。
PHEONA: An Evaluation Framework for Large Language Model-based Approaches to Computational Phenotyping
- 提出专用于医疗表型的评估框架PHEONA
- 在急性呼吸衰竭治疗分类中达到高准确率
- 适合医疗AI研究者和临床数据科学家参考
计算表型在生物医学研究中至关重要,但传统方法依赖大量人工数据审查,耗时耗力。尽管机器学习与自然语言处理技术有所进展,仍需提升。尽管大语言模型(LLMs)在文本任务中表现优异,其在表型任务中的应用研究仍较少。为此,我们开发了面向观察性健康数据表型评估的框架PHEONA,明确了特定场景下的考量因素。我们将PHEONA应用于急性呼吸衰竭(ARF)呼吸支持疗法的概念分类任务。在测试的样本概念中,取得了高分类准确率,表明基于LLM的方法在改善计算表型流程方面具有潜力。
原文摘要 · Abstract (English)
Computational phenotyping is essential for biomedical research but often requires significant time and resources, especially since traditional methods typically involve extensive manual data review. While machine learning and natural language processing advancements have helped, further improvements are needed. Few studies have explored using Large Language Models (LLMs) for these tasks despite known advantages of LLMs for text-based tasks. To facilitate further research in this area, we developed an evaluation framework, Evaluation of PHEnotyping for Observational Health Data (PHEONA), that outlines context-specific considerations. We applied and demonstrated PHEONA on concept classification, a specific task within a broader phenotyping process for Acute Respiratory Failure (ARF) respiratory support therapies. From the sample concepts tested, we achieved high classification accuracy, suggesting the potential for LLM-based methods to improve computational phenotyping processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。