arXiv:2601.12731cs.CLcs.AI2026-01ACL被引 2

发现大模型对难题的判断分两阶段:先抽象共性,再转为语言特有。

A Shared Geometry of Difficulty in Multilingual Language Models

  • 用21种语言的AMC数据训练线性探测器,分析难度信号分布
  • 深层表征准确但跨语言泛化差,浅层表征泛化强但单语性能低
  • 揭示模型从抽象到具体、从通用到语言依赖的推理演化路径

本文研究多语言大模型中问题难度预测的内部表征几何结构。通过在易难基准测试的AMC子集(翻译成21种语言)上训练线性探测器,发现难度信号在模型内部的浅层(早期层)和深层(后期层)分别出现,表现出功能差异。基于深层表示的探测器在同语言任务上表现优异,但跨语言泛化能力差;而基于浅层表示的探测器虽单语性能较低,但在跨语言场景下显著更优。结果表明,大模型首先形成语言无关的难度表征,随后转化为语言特定的表达。这一两阶段机制与现有大模型可解释性研究一致,说明模型先在抽象概念空间中运作,再生成语言特异性输出。该发现表明,这种从通用到具体的表征演进不仅适用于语义内容,也涵盖高阶元认知属性如难度估计。

原文摘要 · Abstract (English)

Predicting problem-difficulty in large language models (LLMs) refers to estimating how difficult a task is according to the model itself, typically by training linear probes on its internal representations. In this work, we study the multilingual geometry of problem-difficulty in LLMs by training linear probes using the AMC subset of the Easy2Hard benchmark, translated into 21 languages. We found that difficulty-related signals emerge at two distinct stages of the model internals, corresponding to shallow (early-layers) and deep (later-layers) internal representations, that exhibit functionally different behaviors. Probes trained on deep representations achieve high accuracy when evaluated on the same language but exhibit poor cross-lingual generalization. In contrast, probes trained on shallow representations generalize substantially better across languages, despite achieving lower within-language performance. Together, these results suggest that LLMs first form a language-agnostic representation of problem difficulty, which subsequently becomes language-specific. This closely aligns with existing findings in LLM interpretability showing that models tend to operate in an abstract conceptual space before producing language-specific outputs. We demonstrate that this two-stage representational process extends beyond semantic content to high-level meta-cognitive properties such as problem-difficulty estimation.

大模型理解多语言可解释性难度预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。