arXiv:2409.10504cs.CL2024-09NeurIPS被引 4

让医学编码模型的预测过程可解释,通过稀疏特征揭示隐藏的医疗概念。

DILA: Dictionary Label Attention for Mechanistic Interpretability in High-dimensional Multi-label Medical Coding Prediction

  • 将密集嵌入分解为稀疏特征,每个非零项对应一个全局学习的医疗概念。
  • 人工评估显示稀疏嵌入可读性比稠密嵌入高至少50%。
  • 用大模型自动识别数千个医疗概念,适合需要透明决策的医疗AI场景。

高维或多标签医学编码预测需要兼具准确性和可解释性。现有方法多依赖局部解释,难以提供多标签预测整体机制的全面说明。我们提出一种机制可解释模块DILA,将难以理解的稠密嵌入分解为稀疏嵌入空间,其中每个非零元素(字典特征)代表一个全局学习的医学概念。人类评估表明,我们的稀疏嵌入在可读性上比稠密嵌入至少提升50%。通过利用大语言模型(LLMs)的自动化字典特征识别流程,我们分析并总结每个字典特征中激活最强的词元,发现了数千个学习到的医学概念。我们通过稀疏可解释矩阵表示字典特征与医学编码之间的关系,增强了对模型预测机制的全局理解,在保持竞争性能和可扩展性的同时,无需大量人工标注。

原文摘要 · Abstract (English)

Predicting high-dimensional or extreme multilabels, such as in medical coding, requires both accuracy and interpretability. Existing works often rely on local interpretability methods, failing to provide comprehensive explanations of the overall mechanism behind each label prediction within a multilabel set. We propose a mechanistic interpretability module called DIctionary Label Attention (\method) that disentangles uninterpretable dense embeddings into a sparse embedding space, where each nonzero element (a dictionary feature) represents a globally learned medical concept. Through human evaluations, we show that our sparse embeddings are more human understandable than its dense counterparts by at least 50 percent. Our automated dictionary feature identification pipeline, leveraging large language models (LLMs), uncovers thousands of learned medical concepts by examining and summarizing the highest activating tokens for each dictionary feature. We represent the relationships between dictionary features and medical codes through a sparse interpretable matrix, enhancing the mechanistic and global understanding of the model's predictions while maintaining competitive performance and scalability without extensive human annotation.

可解释性医学编码稀疏嵌入大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。