用轻量方法提升大模型临床编码准确率,解决近似错误问题
Toward Reliable Clinical Coding with Language Models: Verification and Lightweight Adaptation
- 引入临床编码验证机制,识别层级相近但错误的预测
- 轻量微调与提示工程使准确率显著提升,无需复杂搜索
- 发布新标注数据集,解决现有数据偏倚与证据不全问题
准确的临床编码对医疗记录、计费和决策至关重要。现有研究表明,通用大模型在该任务上表现不佳,而基于精确匹配的评估常忽略层级接近但错误的预测。我们的分析发现,此类层级错配占大模型失败的相当比例。通过提示工程和小规模微调等轻量干预,可在避免搜索类方法计算开销的前提下提升准确性。为应对现有数据集如MIMIC中存在的证据不全与住院患者偏倚问题,我们发布了包含门诊病历与ICD-10编码的专家双标注基准数据集。结果表明,编码验证作为独立任务或流水线环节,是提升大模型医疗编码可靠性的重要步骤。
原文摘要 · Abstract (English)
Accurate clinical coding is essential for healthcare documentation, billing, and decision-making. While prior work shows that off-the-shelf LLMs struggle with this task, evaluations based on exact match metrics often overlook errors where predicted codes are hierarchically close but incorrect. Our analysis reveals that such hierarchical misalignments account for a substantial portion of LLM failures. We show that lightweight interventions, including prompt engineering and small-scale fine-tuning, can improve accuracy without the computational overhead of search-based methods. To address hierarchically near-miss errors, we introduce clinical code verification as both a standalone task and a pipeline component. To mitigate the limitations in existing datasets, such as incomplete evidence and inpatient bias in MIMIC, we release an expert double-annotated benchmark of outpatient clinical notes with ICD-10 codes. Our results highlight verification as an effective and reliable step toward improving LLM-based medical coding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。