arXiv:2605.05741cs.AI2026-05

通过细粒度置信度轨迹量化大模型推理中的认知努力。

HyperLens: Quantifying Cognitive Effort in LLMs with Fine-grained Confidence Trajectory

论文配图:HyperLens: Quantifying Cognitive Effort in LLMs with Fine-grained Confidence Trajectory
图 1 · 摘自论文原文
  • 利用深层注意力机制放大置信度微变,追踪推理过程
  • 复杂任务的置信度轨迹明显发散,对应更高认知努力
  • 揭示苏教微调会降低认知努力并损害性能

尽管大型语言模型(LLMs)在多种任务上表现优异,但其推理动态因现有分析工具分辨率有限而难以理解。本文发现:Transformer架构中深层层会固有放大逐层置信度的小幅变化,从而生成细粒度置信度轨迹。基于此,我们提出HyperLens,一种高分辨率探测器,用于追踪置信度轨迹并量化推理过程中的认知努力。在多个模型与数据集上,HyperLens揭示了复杂任务与简单任务的置信度轨迹存在一致差异。我们将这一模式抽象为可量化的认知努力指标。分析表明:复杂任务始终需要更高的认知努力。最后,我们对标准监督微调(SFT)的常见副作用进行了机制诊断:它可能降低认知努力,进而损害域内任务的表现。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) achieve strong performance across diverse tasks, their inference dynamics remain poorly understood because of the limited resolution of existing analysis tools. In this work, we identify an intrinsic magnification mechanism in transformer architectures: deeper layers inherently magnify the small changes of layer-wise confidence, providing a fine-grained confidence trajectory. Building on this insight, we introduce HyperLens, a high-resolution probe designed to trace confidence trajectories and quantify the cognitive effort during inference. Across LLMs and datasets, HyperLens reveals a consistent divergence in confidence trajectories that separates complex from simple tasks. We abstract this pattern into a quantitative cognitive effort metric. Our analysis reveals a fundamental principle: complex tasks consistently require higher cognitive effort. Finally, we provide a mechanistic diagnosis of a common side effect of standard Supervised Fine-Tuning (SFT): it can reduce cognitive effort and consequently degrade performance on in-domain tasks.

大模型分析认知努力置信度轨迹SFT诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。