arXiv:2509.22888cs.AIcs.CL2025-09中稿 · TMLR被引 4

用几何空间解析大模型能力,让题目难度和模型专长一目了然。

JE-IRT: A Geometric Lens on LLM Abilities through Joint Embedding Item Response Theory

  • 将模型与题目映射到共享几何空间,方向表语义,模长表难度
  • 发现模型在分布外表现由方向对齐决定,模长越大题目越难
  • 可快速添加新模型,还能自动发现模型内部能力分类

主流大模型评估将多元能力压缩为单一分数,掩盖其多维本质。本文提出JE-IRT,一种联合嵌入项目反应理论框架,将大模型与问题共同嵌入同一几何空间。问题嵌入的方向表示语义,模长表示难度,模型对问题的正确性由两者几何交互决定。该几何结构取代全局排名,实现主题专长与相关问题间的平滑过渡。实验表明,分布外行为可通过方向对齐解释,且模长越大问题越难。该空间支持自然泛化:学习后仅需拟合单个嵌入即可加入新模型。空间还揭示了模型内隐的能力分类,与人类定义的主题类别部分重合。此外,简单线性探测即可恢复跨学科能力方向,如识别出病毒学、全球事实等看似无关领域中具有量化挑战性的题目。JE-IRT建立了一个统一且可解释的几何视角,连接模型能力与题目结构,为模型评估与泛化提供新思路。

原文摘要 · Abstract (English)

Standard LLM evaluation practices compress diverse abilities into single scores, obscuring their inherently multidimensional nature. We present JE-IRT, a geometric item-response framework that embeds both LLMs and questions in a shared space. For question embeddings, the direction encodes semantics and the norm encodes difficulty, while correctness on each question is determined by the geometric interaction between the model and question embeddings. This geometry replaces a global ranking of LLMs with topical specialization and enables smooth variation across related questions. Building on this framework, our experimental results reveal that out-of-distribution behavior can be explained through directional alignment, and that larger norms consistently indicate harder questions. Moreover, JE-IRT naturally supports generalization: once the space is learned, new LLMs are added by fitting a single embedding. The learned space further reveals an LLM-internal taxonomy that only partially aligns with human-defined subject categories. We also show that simple linear probes of the embedding space recover cross-subject ability directions, such as an arithmetic axis that highlights quantitatively demanding questions in seemingly distant subjects like virology and global facts. JE-IRT thus establishes a unified and interpretable geometric lens that connects LLM abilities with the structure of questions, offering a distinctive perspective on model evaluation and generalization.

大模型评估几何嵌入能力分析项目反应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。