arXiv:2604.03589cs.AI2026-04中稿 · Publish it in 12th…

分析小模型在问答任务中熵与注意力的动态,揭示其幻觉成因。

Entropy and Attention Dynamics in Small Language Models: A Trace-Level Structural Analysis on the TruthfulQA Benchmark

  • 通过逐令牌分析熵与注意力分布,揭示模型内部不确定性演化机制。
  • 发现四类模型呈现三类熵模式:确定型、探索型、平衡型,对应不同输出稳定性。
  • 结果可指导边缘设备小模型设计,提升事实准确性与可靠性。

小语言模型(SLMs)因其轻量特性被广泛部署于资源受限场景,但其常产生自信的错误预测和不稳定的输出,难以胜任事实性与决策关键任务。现有评估方法仅关注最终准确率或幻觉率,忽视内部行为对输出的影响。本文针对TruthfulQA数据集,对参数量为1B-1.7B的四款模型进行逐层追踪分析,考察解码过程中输出熵、注意力熵、头分散度及隐藏状态表示的动态变化。结果显示,模型可按熵演化模式分为三类:确定型(DeepSeek-1.5B、LLaMA-1B)——熵随时间下降;探索型(Gemma-1B)——熵持续上升;平衡型(Qwen-1.7B)——熵稳定且适中。三类模型在隐藏状态迁移与注意力分布上亦表现出显著差异。研究证明,模型真实性源于结构化的熵与注意力动态。监测并优化这些内部不确定性模式,有助于设计更可靠、抗幻觉、适配特定场景的边缘小模型。

原文摘要 · Abstract (English)

Small language models (SLMs) have been increasingly deployed in edge devices and other resource-constrained settings. However, these models make confident mispredictions and produce unstable output, making them risky for factual and decision-critical tasks. Current evaluation methodology relies on final accuracy or hallucination rates without explaining how internal model behavior affects outputs. Specifically, how entropy evolves during decoding, how attention is distributed across layers, and how hidden representations contribute to uncertainty, logical inconsistencies, and misinformation propagation are often overlooked. Consequently, this study introduces a trace-level analysis of entropy and attention dynamics in SLMs evaluated with the TruthfulQA dataset. Four models with parameter ranges of 1B-1.7B parameters were examined via token-level output entropy, attention entropy, head dispersion, and hidden-state representation. The results reflect three model classifications by entropy patterns. Deterministic models (DeepSeek-1.5B and LLaMA-1B): output entropy decreases over time. Exploratory models (Gemma-1B): with increasing entropy, and balanced models (Qwen-1.7B): have moderate and stable entropy. Also, each group has distinctively different hidden-state movement and attention dispersion patterns. The analysis demonstrates that truthfulness in SLMs emerges from structured entropy and attention dynamics. Monitoring and optimizing these internal uncertainty patterns can guide the design of a more reliable, hallucination-aware, and application-specific edge SLMs.

小模型熵分析注意力机制幻觉检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。