arXiv:2608.20887cs.CLcs.AI2026-08

用知识引导大模型精准生成医疗编码,提升合规性与准确性

KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs

论文配图:KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs
图 1 · 摘自论文原文
  • 将临床证据与官方编码规则结合,让大模型推理更可靠
  • 在多个基准数据集上超越现有主流方法,平均提升1.8个点
  • 适合医疗信息管理、医保自动化等需要高准确编码的场景

自动医疗编码(AMC)是为临床记录分配标准化国际疾病分类(ICD)代码的关键任务,对医疗报销、质量评估和临床研究至关重要。现有基于预训练语言模型的方法通常将AMC视为固定代码集上的极端多标签分类问题,而近期基于大语言模型(LLM)的方法则将其转化为生成或多步推理任务。然而,仍面临临床文本过长难以理解、ICD标签空间庞大以及编码规则复杂且未被LLM显式建模等挑战。本文提出知识引导的临床证据推理框架KREL,利用LLM理解临床文本,并整合外部ICD编码指南作为结构化知识,实现领域知识与模型推理的紧密耦合,降低幻觉并提升编码标准符合度。在多个基准数据集上的实验表明,KREL持续优于强基线的PLM方法及先进的LLM方法。

原文摘要 · Abstract (English)

Automatic Medical Coding (AMC), which assigns standardized International Classification of Diseases (ICD) codes to clinical notes, is essential for medical reimbursement, quality reporting, and clinical research. Existing pre-trained language model (PLM)-based methods typically formulate AMC as an extreme multi-label classification problem over a predefined code set, while recent large language model (LLM)-based approaches instead frame it as generation or multi-step reasoning. However, key challenges remain, including the extreme length of clinical notes that hinders effective interpretation, the vast ICD label space, and complex coding rules that are not explicitly captured by LLMs. In this work, we propose Knowledge-Guided Reasoning over Clinical Evidence with LLMs (KREL), a framework that leverages LLMs for clinical text understanding and reasoning while integrating external ICD coding guidelines as structured knowledge. This design enables tight coupling between domain knowledge and LLM reasoning, reducing hallucinations and improving compliance with coding standards. Experiments on benchmark datasets show that KREL consistently outperforms strong PLM-based and state-of-the-art LLM-based baselines.

医疗编码大模型知识融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。