arXiv:2411.00173cs.CLcs.AI2024-11EMNLP被引 7

用词典学习提升医学编码模型可解释性,让无关词也能被合理解读。

Beyond Label Attention: Transparency in Language Models for Automated Medical Coding via Dictionary Learning

  • 通过词典学习从模型嵌入中提取稀疏激活特征,替代传统注意力机制。
  • 90%以上无关临床词的隐藏含义被揭示,解释更符合医学逻辑。
  • 结果可被人类理解,适合需要高可信度医疗AI的场景。

医学编码是将非结构化临床文本转化为标准化医疗代码的关键但耗时的医疗实践。尽管大语言模型(LLM)可自动化该过程并提高效率,但可解释性对维护患者信任至关重要。当前医疗编码应用的可解释性主要依赖标签注意力机制,常导致无关词被错误突出。本文提出基于词典学习的方法,能从密集语言模型嵌入中高效提取稀疏激活表示。相比常规标签注意力,本方法构建可解释词典,增强对每个ICD代码预测的机制性解释,即使被突出的词在医学上无关。实验表明,词典特征可引导模型行为,阐明超过90%医学无关词的隐藏意义,且具备人类可读性。

原文摘要 · Abstract (English)

Medical coding, the translation of unstructured clinical text into standardized medical codes, is a crucial but time-consuming healthcare practice. Though large language models (LLM) could automate the coding process and improve the efficiency of such tasks, interpretability remains paramount for maintaining patient trust. Current efforts in interpretability of medical coding applications rely heavily on label attention mechanisms, which often leads to the highlighting of extraneous tokens irrelevant to the ICD code. To facilitate accurate interpretability in medical language models, this paper leverages dictionary learning that can efficiently extract sparsely activated representations from dense language model embeddings in superposition. Compared with common label attention mechanisms, our model goes beyond token-level representations by building an interpretable dictionary which enhances the mechanistic-based explanations for each ICD code prediction, even when the highlighted tokens are medically irrelevant. We show that dictionary features can steer model behavior, elucidate the hidden meanings of upwards of 90% of medically irrelevant tokens, and are human interpretable.

医学编码可解释性词典学习LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。