arXiv:2506.10378cs.LGcs.AI2025-06被引 5

用因果模型发现大模型能力的层级结构。

Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning

  • 通过线性因果结构建模隐藏能力因子,控制基础模型干扰。
  • 发现三大能力:通用解题→指令遵循→数学推理,有明确因果链。
  • 适合关注模型能力本质、评估方法改进的研究者。

语言模型能力的可靠评估对指导模型发展至关重要。然而,该领域的严格因果评估面临重大方法挑战,包括复杂的混杂效应以及大量重训练带来的高昂计算成本。为此,我们提出一种因果表示学习框架,将观测到的基准性能建模为少数潜在能力因子的线性变换。关键在于,在适当控制基础模型作为共同混杂因子后,这些潜在因子被识别为因果相关。将该方法应用于涵盖超过1500个模型、来自Open LLM Leaderboard六个基准的综合性数据集,我们识别出一个可靠的三节点线性因果结构,可解释观测到的性能变化。进一步解读该因果结构揭示了超越简单排名的科学洞见:从通用问题求解能力出发,经由指令遵循能力,最终导向数学推理能力,存在清晰的因果方向。结果强调了在评估中严谨控制基础模型差异的重要性,这是准确揭示潜在模型能力之间因果关系的关键步骤。

原文摘要 · Abstract (English)

Faithful evaluation of language model capabilities is crucial for deriving actionable insights that can inform model development. However, rigorous causal evaluations in this domain face significant methodological challenges, including complex confounding effects and prohibitive computational costs associated with extensive retraining. To tackle these challenges, we propose a causal representation learning framework wherein observed benchmark performance is modeled as a linear transformation of a few latent capability factors. Crucially, these latent factors are identified as causally interrelated after appropriately controlling for the base model as a common confounder. Applying this approach to a comprehensive dataset encompassing over 1500 models evaluated across six benchmarks from the Open LLM Leaderboard, we identify a concise three-node linear causal structure that reliably explains the observed performance variations. Further interpretation of this causal structure provides substantial scientific insights beyond simple numerical rankings: specifically, we reveal a clear causal direction starting from general problem-solving capabilities, advancing through instruction-following proficiency, and culminating in mathematical reasoning ability. Our results underscore the essential role of carefully controlling base model variations during evaluation, a step critical to accurately uncovering the underlying causal relationships among latent model capabilities.

大模型能力因果推理能力分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。