用AI自动补全临床试验中的医学编码,让老数据更好用。
Unlocking Historical Clinical Trial Data with ALIGN: A Compositional Large Language Model System for Medical Coding
- 分三步生成、自评并打分医学编码,支持零样本推理。
- 对罕见药物编码准确率提升17%,整体准确率达90%。
- 成本极低,适合医院和药企快速部署使用。
历史临床试验数据的再利用对加速医学研究和药物开发具有重要意义,但因缺乏标准医学编码导致数据难以互通。现有大语言模型在复杂编码任务上表现受限。本文提出ALIGN系统,通过三步流程实现零样本医学编码:生成多样候选编码、自评估编码质量、输出置信度并支持人工介入。在22个免疫学试验数据上测试,将药物术语映射至ATC编码、病史术语映射至MedDRA编码。对于MedDRA编码,各层级准确率高,尤其在最高细分层级(HLGT)达87-90%;对于ATC编码,整体准确率达72-73%,常见药物达86-89%,优于基线7-22%。基于不确定性拒答机制使准确率提升至90%,仅需30%人工介入。每条编码成本低至0.0007美元(GPT-4o-mini)至0.02美元(GPT-4o),显著降低临床应用门槛。系统有效提升临床数据可互操作性与复用性。
原文摘要 · Abstract (English)
The reuse of historical clinical trial data has significant potential to accelerate medical research and drug development. However, interoperability challenges, particularly with missing medical codes, hinders effective data integration across studies. While Large Language Models (LLMs) offer a promising solution for automated coding without labeled data, current approaches face challenges on complex coding tasks. We introduce ALIGN, a novel compositional LLM-based system for automated, zero-shot medical coding. ALIGN follows a three-step process: (1) diverse candidate code generation; (2) self-evaluation of codes and (3) confidence scoring and uncertainty estimation enabling human deferral to ensure reliability. We evaluate ALIGN on harmonizing medication terms into Anatomical Therapeutic Chemical (ATC) and medical history terms into Medical Dictionary for Regulatory Activities (MedDRA) codes extracted from 22 immunology trials. ALIGN outperformed the LLM baselines, while also providing capabilities for trustworthy deployment. For MedDRA coding, ALIGN achieved high accuracy across all levels, matching RAG and excelling at the most specific levels (87-90% for HLGT). For ATC coding, ALIGN demonstrated superior performance, particularly at lower hierarchy levels (ATC Level 4), with 72-73% overall accuracy and 86-89% accuracy for common medications, outperforming baselines by 7-22%. ALIGN's uncertainty-based deferral improved accuracy by 17% to 90% accuracy with 30% deferral, notably enhancing performance on uncommon medications. ALIGN achieves this cost-efficiently at \$0.0007 and \$0.02 per code for GPT-4o-mini and GPT-4o, reducing barriers to clinical adoption. ALIGN advances automated medical coding for clinical trial data, contributing to enhanced data interoperability and reusability, positioning it as a promising tool to improve clinical research and accelerate drug development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。