研究注意力可视化对医学文献分类解释力的影响
User Perception of Attention Visualizations: Effects on Interpretability Across Evidence-Based Medical Documents
- 用用户实验测试不同注意力可视化方式在医学文献分类中的效果
- 医生认为亮度/背景色比条形长度更易理解,与经典视觉原则相反
- 注意力权重本身解释力有限,但呈现方式显著影响使用感受
注意力机制是Transformer架构的核心组件。除了提升性能,注意力权重也被视为可解释性工具,其数值大小可能反映输入特征(如文档中的词元)与模型预测的相关性。在循证医学中,此类解释有助于医生理解并交互于用于分类生物医学文献的AI系统。然而,目前尚无共识认为注意力权重能提供有效解释,且极少研究探讨注意力可视化对其解释效用的影响。为此,我们开展了一项用户研究,评估注意力解释是否有助于医学专家进行生物医学文献分类,并探索最优可视化方式。研究对象为来自不同领域的医学专家,任务是根据研究设计(如系统综述、广义综述、随机与非随机试验)分类文章。结果显示,基于XLNet的Transformer模型分类准确;但注意力权重未被普遍视为有用的解释依据。然而,这种感知差异显著依赖于可视化形式。与Munzner的视觉有效性原则(偏好精确编码如条形长度)相反,用户更倾向使用文本亮度或背景颜色等直观形式。尽管结果未证实注意力权重的整体解释价值,但表明其感知帮助程度受呈现方式显著影响。
原文摘要 · Abstract (English)
The attention mechanism is a core component of the Transformer architecture. Beyond improving performance, attention has been proposed as a mechanism for explainability via attention weights, which are associated with input features (e.g., tokens in a document). In this context, larger attention weights may imply more relevant features for the model's prediction. In evidence-based medicine, such explanations could support physicians' understanding and interaction with AI systems used to categorize biomedical literature. However, there is still no consensus on whether attention weights provide helpful explanations. Moreover, little research has explored how visualizing attention affects its usefulness as an explanation aid. To bridge this gap, we conducted a user study to evaluate whether attention-based explanations support users in biomedical document classification and whether there is a preferred way to visualize them. The study involved medical experts from various disciplines who classified articles based on study design (e.g., systematic reviews, broad synthesis, randomized and non-randomized trials). Our findings show that the Transformer model (XLNet) classified documents accurately; however, the attention weights were not perceived as particularly helpful for explaining the predictions. However, this perception varied significantly depending on how attention was visualized. Contrary to Munzner's principle of visual effectiveness, which favors precise encodings like bar length, users preferred more intuitive formats, such as text brightness or background color. While our results do not confirm the overall utility of attention weights for explanation, they suggest that their perceived helpfulness is influenced by how they are visually presented.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。