arXiv:2603.29709cs.AIcs.LG2026-03

让机器像医生一样精准理解病历并解释编码依据

Symphony for Medical Coding: A Next-Generation Agentic System for Scalable and Explainable Medical Coding

  • 通过直接读取编码指南,模拟人类医生的推理过程
  • 在多个真实医疗场景中达到当前最佳性能,支持跨系统使用
  • 可定位每条编码对应的原文片段,提升诊断可信度

医学编码将自由文本的临床记录转化为分类系统中的标准化代码,涵盖数万条目且每年更新。它在医保结算、临床研究和质量报告中至关重要,但目前仍高度依赖人工,效率低且易出错。现有自动化方法仅能预测固定代码集,无法适应新代码或不同编码体系,且缺乏预测解释,难以在安全敏感场景中建立信任。本文提出Symphony医学编码系统,模仿资深编码员的工作方式:基于临床叙述与编码指南进行推理。该设计使系统可适配任意编码体系,并提供细粒度证据,明确标注每个代码所对应的原文片段。我们在两个公开基准和三个真实世界数据集(涵盖美国及英国的住院、门诊、急诊及亚专科场景)上评估,结果表明Symphony在所有场景下均达到领先水平,具备灵活部署能力,可作为下一代自动医学编码的基础框架。

原文摘要 · Abstract (English)

Medical coding translates free-text clinical documentation into standardized codes drawn from classification systems that contain tens of thousands of entries and are updated annually. It is central to billing, clinical research, and quality reporting, yet remains largely manual, slow, and error-prone. Existing automated approaches learn to predict a fixed set of codes from labeled data, thereby preventing adaptation to new codes or different coding systems without retraining on different data. They also provide no explanation for their predictions, limiting trust in safety-critical settings. We introduce Symphony for Medical Coding, a system that approaches the task the way expert human coders do: by reasoning over the clinical narrative with direct access to the coding guidelines. This design allows Symphony to operate across any coding system and to provide span-level evidence linking each predicted code to the text that supports it. We evaluate on two public benchmarks and three real-world datasets spanning inpatient, outpatient, emergency, and subspecialty settings across the United States and the United Kingdom. Symphony achieves state-of-the-art results across all settings, establishing itself as a flexible, deployment-ready foundation for automated clinical coding.

医学编码可解释性智能代理临床AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。