arXiv:2411.10069cs.CLcs.PF2024-11被引 6

通过激活方差-稀疏性分析,识别大模型中冗余层并降低幻觉生成。

Layer Importance and Hallucination Analysis in Large Language Models via Enhanced Activation Variance-Sparsity

  • 用激活方差与稀疏性结合评分,量化各层重要性。
  • 剪掉25%低重要层后性能仍保留超90%。
  • 新指标可定位幻觉高发层,提升模型可靠性。

评估大语言模型(LLM)中不同层的重要性对优化性能和提升可解释性至关重要。本文首先提出激活方差-稀疏性得分(AVSS),通过归一化激活方差与稀疏性综合量化每层贡献。基于AVSS排名并剪枝最低影响的25%层,在问答、语言建模和情感分类任务上,性能保持超过90%,揭示了LLM架构中的潜在冗余。在此基础上,我们提出改进版方法EAVSS,引入针对幻觉的激活方差(HSAV)与稀疏性(HSS)指标,精准识别幻觉高发层。通过在这些层上引入对比学习,有效抑制幻觉生成,最大性能提升达12%。在NQ、SciQ、TriviaQA、TruthfulQA和WikiQA数据集上的实验验证了该方法的有效性,为层重要性评估与幻觉缓解提供了完整框架。

原文摘要 · Abstract (English)

Evaluating the importance of different layers in large language models (LLMs) is crucial for optimizing model performance and interpretability. This paper first explores layer importance using the Activation Variance-Sparsity Score (AVSS), which combines normalized activation variance and sparsity to quantify each layer's contribution to overall model performance. By ranking layers based on AVSS and pruning the least impactful 25\%, our experiments on tasks such as question answering, language modeling, and sentiment classification show that over 90\% of the original performance is retained, highlighting potential redundancies in LLM architectures. Building on AVSS, we propose an enhanced version tailored to assess hallucination propensity across layers (EAVSS). This improved approach introduces Hallucination-Specific Activation Variance (HSAV) and Hallucination-Specific Sparsity (HSS) metrics, allowing precise identification of hallucination-prone layers. By incorporating contrastive learning on these layers, we effectively mitigate hallucination generation, contributing to more robust and efficient LLMs(The maximum performance improvement is 12\%). Our results on the NQ, SciQ, TriviaQA, TruthfulQA, and WikiQA datasets demonstrate the efficacy of this method, offering a comprehensive framework for both layer importance evaluation and hallucination mitigation in LLMs.

大模型分析幻觉检测层重要性稀疏性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。