arXiv:2509.18846cs.AI2025-09

用大模型当裁判选最佳编码模型,提升医疗诊断编码准确率。

Model selection meets clinical semantics: Optimizing ICD-10-CM prediction via LLM-as-Judge evaluation, redundancy-aware sampling, and section-aware fine-tuning

  • 以大模型评分筛选最优基础模型,结合语义去重采样数据
  • 在两个医院数据集上,性能超越基线模型,多包含临床章节效果更优
  • 适合医疗信息化、AI辅助诊断系统开发者参考落地

准确的国际疾病分类(ICD)编码对临床记录、医保报销和医疗分析至关重要,但目前仍依赖人工且易出错。尽管大语言模型(LLMs)在自动化编码方面展现潜力,其基础模型选择、输入上下文理解及训练数据冗余等问题限制了实际效果。本文提出一个模块化框架,用于 ICD-10-CM 编码预测,通过有原则的模型选择、去冗余数据采样和结构化输入设计解决上述问题。该框架采用 LLM-as-judge 评估机制与 Plackett-Luce 聚合方法,基于对 ICD-10-CM 编码定义的理解能力,评估并排序开源大模型。引入基于嵌入的相似性度量,实施去冗余采样策略以去除语义重复的出院记录。利用台湾多家医院的结构化出院记录,评估上下文影响,并比较通用与分章节建模范式下的内容覆盖情况。在两个机构数据集上的实验表明,经过微调后选定的模型在内部和外部评估中均持续优于基线模型。更多临床章节信息的引入可显著提升预测性能。本研究使用开源大模型,建立了一套可实践、有依据的 ICD-10-CM 编码预测方法,融合明智的模型选择、高效数据优化与上下文感知提示,为真实世界医疗编码系统部署提供可扩展、机构可用的解决方案。

原文摘要 · Abstract (English)

Accurate International Classification of Diseases (ICD) coding is critical for clinical documentation, billing, and healthcare analytics, yet it remains a labour-intensive and error-prone task. Although large language models (LLMs) show promise in automating ICD coding, their challenges in base model selection, input contextualization, and training data redundancy limit their effectiveness. We propose a modular framework for ICD-10 Clinical Modification (ICD-10-CM) code prediction that addresses these challenges through principled model selection, redundancy-aware data sampling, and structured input design. The framework integrates an LLM-as-judge evaluation protocol with Plackett-Luce aggregation to assess and rank open-source LLMs based on their intrinsic comprehension of ICD-10-CM code definitions. We introduced embedding-based similarity measures, a redundancy-aware sampling strategy to remove semantically duplicated discharge summaries. We leverage structured discharge summaries from Taiwanese hospitals to evaluate contextual effects and examine section-wise content inclusion under universal and section-specific modelling paradigms. Experiments across two institutional datasets demonstrate that the selected base model after fine-tuning consistently outperforms baseline LLMs in internal and external evaluations. Incorporating more clinical sections consistently improves prediction performance. This study uses open-source LLMs to establish a practical and principled approach to ICD-10-CM code prediction. The proposed framework provides a scalable, institution-ready solution for real-world deployment of automated medical coding systems by combining informed model selection, efficient data refinement, and context-aware prompting.

医疗编码大模型评估数据去重临床自然语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。