arXiv:2605.30018cs.CLcs.LG2026-05

用隐层特征分析大模型真实能力,超越表面得分。

Latent Performance Profiling of Large Language Models

论文配图:Latent Performance Profiling of Large Language Models
图 1 · 摘自论文原文
  • 通过隐层激活和输出分布提取通用诊断指标
  • 相同基准分数的模型可能有截然不同的内在表现
  • 适合关注模型可靠性与安全性的研究者

大语言模型(LLMs)在标准化评测中常取得优异成绩,但仅靠准确率无法全面反映其能力。现有开源模型评测存在数据污染、任务范围狭窄、与真实可靠性对齐不足等问题。如MMLU PRO、BBH或IFEval等基于基准的评估主要捕捉模型在固定测试集上的输出,而非其信息处理、置信度校准或内部知识结构。本文倡导从基准主导评估转向以状态为中心的内在评估,提出隐性能剖析(Latent Performance Profiling, LPP)框架——从隐藏激活和输出分布中提取与任务无关的诊断指标。LPP在模型隐层表示与动态中定义一组标量指标,揭示可跨规模比较的独立于尺度的特性,提供稳定且敏感于架构的表征。在8个规模0.5B-14B的LLM上进行实证分析表明,相似基准得分的模型可能表现出显著不同的隐层特征,如熵值或适应性差异。基于这些洞察,我们设计了与内在指标对齐的合成探测器,用于不确定性与符号推理评估,同时规避排行榜偏见。建议在报告基准分数的同时引入LPP,以获得更深入、可解释的模型行为理解,支持更可靠的模型选择、安全评估及超越表面准确率的评测。

原文摘要 · Abstract (English)

Large language models (LLMs) frequently achieve impressive scores on standardized benchmarks, yet accuracy alone offers a limited view of their capabilities. Evaluating open-source LLMs through leaderboards faces persistent issues like data contamination, narrow task scope, and weak alignment with real-world reliability. Benchmark-based evaluations such as MMLU PRO, BBH, or IFEval primarily capture what a model outputs on fixed test sets, not how it processes information, calibrates uncertainty, or structures internal knowledge. In this article, we advocate for a shift from benchmark-centric evaluation toward a complementary, state-centered intrinsic assessment of LLMs. To this end, we introduce Latent Performance Profiling (LPP) -- a framework that derives task-agnostic diagnostics from hidden activations and output distributions. LPP defines a set of scalar metrics on a model's latent representations and dynamics, revealing scale-independent traits that enable interpretable comparisons and uncover hidden vulnerabilities. Unlike static accuracy scores, LPP provides stable, architecture-sensitive signatures across models of similar size. With extensive empirical analyses across eight LLMs, spanning a size range of 0.5B-14B, we demonstrate that models with similar benchmark scores can exhibit contrasting latent profiles, such as differences in entropy or adaptability. Guided by these insights, we design synthetic probes for uncertainty and symbolic reasoning that align with intrinsic metrics while decoupling from leaderboard bias. We recommend that reporting LPP alongside benchmarks provides a deeper, interpretable understanding of model behavior, enabling more reliable model selection, safety assessment, and evaluation beyond surface-level accuracy.

大模型评估隐层分析模型可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。