arXiv:2501.05479cs.CLcs.LG2025-01被引 3

小模型微调后在手术编码上表现媲美大模型,兼顾准确与隐私。

Practical Design and Benchmarking of Generative AI Applications for Surgical Billing and Coding

  • 用机构数据微调小模型,提升编码准确性。
  • 微调后模型对ICD-10和CPT编码召回率超72%,伪造率低于1%。
  • 适合资源有限但需高隐私保护的医疗AI应用。

背景:医疗领域存在大量可由生成式人工智能(AI)自动化或增强的流程,如医疗计费与编码。然而,现有基础大语言模型在生成国际疾病分类第10版临床修改版(ICD-10-CM)和当前程序术语(CPT)代码时表现不佳。此外,生成式AI在医疗应用中还面临安全与财务挑战。本文提出一种面向医疗计费与编码的生成式AI工具开发策略,平衡准确性、可及性与患者隐私。方法:使用机构数据微调Phi-3 Mini和Phi-3 Medium模型,并与Phi-3基线模型、Phi-3 RAG应用及GPT-4o对比。输入为术后手术报告,目标输出为关联的医保账单,包含ICD-10、CPT及修饰符代码。性能指标包括代码生成准确率、无效代码比例及账单格式保真度。结果:两个微调模型表现优于或等同于GPT-4o。Phi-3 Medium微调模型表现最佳(ICD-10召回率与精确率:72%、72%;CPT召回率与精确率:77%、79%;修饰符召回率与精确率:63%、64%)。该模型仅伪造1%的ICD-10代码和0.6%的CPT代码。结论:仅用少量开源工具和低投入,对特定任务进行领域数据微调的小模型,可达到与大型主流消费级模型相当的性能。

原文摘要 · Abstract (English)

Background: Healthcare has many manual processes that can benefit from automation and augmentation with Generative Artificial Intelligence (AI), the medical billing and coding process. However, current foundational Large Language Models (LLMs) perform poorly when tasked with generating accurate International Classification of Diseases, 10th edition, Clinical Modification (ICD-10-CM) and Current Procedural Terminology (CPT) codes. Additionally, there are many security and financial challenges in the application of generative AI to healthcare. We present a strategy for developing generative AI tools in healthcare, specifically for medical billing and coding, that balances accuracy, accessibility, and patient privacy. Methods: We fine tune the PHI-3 Mini and PHI-3 Medium LLMs using institutional data and compare the results against the PHI-3 base model, a PHI-3 RAG application, and GPT-4o. We use the post operative surgical report as input and the patients billing claim the associated ICD-10, CPT, and Modifier codes as the target result. Performance is measured by accuracy of code generation, proportion of invalid codes, and the fidelity of the billing claim format. Results: Both fine-tuned models performed better or as well as GPT-4o. The Phi-3 Medium fine-tuned model showed the best performance (ICD-10 Recall and Precision: 72%, 72%; CPT Recall and Precision: 77%, 79%; Modifier Recall and Precision: 63%, 64%). The Phi-3 Medium fine-tuned model only fabricated 1% of ICD-10 codes and 0.6% of CPT codes generated. Conclusions: Our study shows that a small model that is fine-tuned on domain-specific data for specific tasks using a simple set of open-source tools and minimal technological and monetary requirements performs as well as the larger contemporary consumer models.

医疗AI编码生成小模型微调隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。