首次系统分析S4D模型在代码漏洞检测中的长程依赖理解能力
Analysis of Long Range Dependency Understanding in State Space Models
- 通过时域频域分析S4D核函数,揭示其滤波特性变化机制
- 不同架构下S4D核可表现为低通、带通或高通滤波器,影响建模效果
- 为设计更优S4D模型提供可解释性指导,适合模型优化研究者
尽管状态空间模型(SSMs)在长序列基准上表现优异,但多数研究聚焦预测准确率而忽视可解释性。本文首次对在真实任务(源码漏洞检测)上训练的对角化状态空间模型(S4D)进行系统的核可解释性研究。通过时域与频域分析S4D核,发现其长程建模能力随模型架构显著变化。例如,不同架构下S4D核可呈现低通、带通或高通滤波特性。这些发现为未来设计更优S4D模型提供了重要指导。
原文摘要 · Abstract (English)
Although state-space models (SSMs) have demonstrated strong performance on long-sequence benchmarks, most research has emphasized predictive accuracy rather than interpretability. In this work, we present the first systematic kernel interpretability study of the diagonalized state-space model (S4D) trained on a real-world task (vulnerability detection in source code). Through time and frequency domain analysis of the S4D kernel, we show that the long-range modeling capability of S4D varies significantly under different model architectures, affecting model performance. For instance, we show that the depending on the architecture, S4D kernel can behave as low-pass, band-pass or high-pass filter. The insights from our analysis can guide future work in designing better S4D-based models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。