arXiv:2508.16777cs.AI2025-08Conference of the …被引 1

用新数据集评估并提升医疗编码理由的可信与合理,让AI解释更可懂。

Evaluation and LLM-Guided Learning of ICD Coding Rationales

  • 构建多粒度医疗编码理由数据集,支持系统评估
  • 大模型生成理由在人类评估中表现最优,可信度高
  • 利用大模型理由做远监督,提升模型自动生成能力

ICD编码是将电子病历中的非结构化文本映射到国际疾病分类标准代码的过程。为增强模型可信度与透明性,现有研究多依赖注意力机制生成理由和医生定性评估,但缺乏统一标准和高质量标注数据集来系统评估多种理由类型。本文聚焦理由的忠实性与合理性两个核心维度,构建基于MIMIC-IV数据库和ICD-10系统的新型多粒度理由标注数据集。评估三种理由类型:实体链接生成的提及、大模型生成的理由、以及编码模型注意力得分。发现大模型生成理由在人类评估中合理性突出。进一步利用其作为远监督信号,训练出可生成更合理理由的新方法,并通过少量人工标注示例微调大模型,显著提升教师与学生模型生成理由的合理性。

原文摘要 · Abstract (English)

ICD coding is the process of mapping unstructured text from Electronic Health Records (EHRs) to standardised codes defined by the International Classification of Diseases (ICD) system. In order to promote trust and transparency, existing explorations on the explainability of ICD coding models primarily rely on attention-based rationales and qualitative assessments conducted by physicians, yet lack a systematic evaluation across diverse types of rationales using consistent criteria and high-quality rationale-annotated datasets specifically designed for the ICD coding task. Moreover, dedicated methods explicitly trained to generate plausible rationales remain scarce. In this work, we present evaluations of the explainability of rationales in ICD coding, focusing on two fundamental dimensions: faithfulness and plausibility -- in short how rationales influence model decisions and how convincing humans find them. For plausibility, we construct a novel, multi-granular rationale-annotated ICD coding dataset, based on the MIMIC-IV database and the updated ICD-10 coding system. We conduct a comprehensive evaluation across three types of ICD coding rationales: entity-level mentions automatically constructed via entity linking, LLM-generated rationales, and rationales based on attention scores of ICD coding models. Building upon the strong plausibility exhibited by LLM-generated rationales, we further leverage them as distant supervision signals to develop rationale learning methods. Additionally, by prompting the LLM with few-shot human-annotated examples from our dataset, we achieve notable improvements in the plausibility of rationale generation in both the teacher LLM and the student rationale learning models.

医疗AI可解释性大模型应用自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。