arXiv:2608.13048cs.AI2026-08

用动态模态分解解析大模型分类决策,让预测过程可解释。

DMDIntel: Interpreting Large Language Models via Dynamic Mode Decomposition

论文配图:DMDIntel: Interpreting Large Language Models via Dynamic Mode Decomposition
图 1 · 摘自论文原文
  • 通过动态模态分解提取大模型隐藏状态中的关键模式
  • 在三个数据集上对输入词元排序的准确率显著优于现有方法
  • 适合关注大模型决策透明性的研究人员和开发者

本文提出DMDIntel,利用动态模态分解(DMD)使大语言模型在分类任务中的预测结果可解释。该方法构建了一个输入归因流程:首先将大模型的隐藏状态分解为显著模式(即模态),然后根据输入词元在这些模态上的投影值进行排序。在三个数据集和三种模型家族上的严格实验表明,DMDIntel获得的输入词元归因排名远超主成分分析、积分梯度和SHAP等当前最优技术。

原文摘要 · Abstract (English)

In this work, we introduce DMDIntel which uses dynamic mode decomposition (DMD) to make the predictions made by LLMs in a classification task interpretable. It develops an input attribution pipeline, that first decomposes the hidden states of an LLM into prominent patterns, also known as modes, and then associates ranks to the input tokens based on the projection values on those modes. Rigorous experiments across three datasets and three model families consistently show that the ranked attribution of input tokens obtained using DMDIntel by far outperforms state-of-the-art techniques such as principal component analysis, integrated gradients and SHAP.

可解释性大模型动态模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。