arXiv:2605.15436cs.CLcs.LG2026-05中稿 · IEEE BigData 2025

对比6种大模型在12类认知任务中的神经激活模式,发现数学推理注意力最混乱,解码器更稀疏。

Neural Activation Patterns Across Language Model Architectures: A Comprehensive Analysis of Cognitive Task Performance

  • 测量激活值、注意力熵和稀疏性,比较六种模型架构的神经模式
  • 数学推理任务注意力熵最高,解码器模型稀疏性显著高于编码器
  • 揭示模型任务特异性行为,指导大模型选型与优化

本文对六种不同大型语言模型(LLM)架构在十二类认知任务上的神经激活模式进行了全面分析。通过系统测量最终激活值、注意力熵和稀疏性模式,揭示了编码器与解码器架构在处理多样化认知任务时的根本差异。对144个任务-模型组合的分析表明,数学推理任务在所有架构中均产生最高的注意力熵,而解码器模型的稀疏性模式显著高于编码器模型。研究结果为现代语言模型的计算特性及其任务特异性神经行为提供了关键洞见,对大数据应用中的模型选择与优化具有重要启示。

原文摘要 · Abstract (English)

This paper presents a comprehensive analysis of neural activation patterns across six distinct large language model (LLM) architectures, examining their performance on twelve cognitive task categories. Through systematic measurement of final activation values, attention entropy, and sparsity patterns, we reveal fundamental differences in how encoder and decoder architectures process diverse cognitive tasks. Our analysis of 144 task-model combinations demonstrates that mathematical reasoning consistently produces the highest attention entropy across all architectures, while decoder models exhibit significantly higher sparsity patterns compared to encoder models. The findings provide critical insights into the computational characteristics of modern language models and their task-specific neural behaviors, with implications for model selection and optimization in big data applications.

语言模型认知任务注意力熵稀疏性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。