arXiv:2510.00844cs.AI2025-10被引 9

用心理测量学方法压缩大模型能力表示,提升模型路由与预测效率

Learning Compact Representations of LLM Abilities via Item Response Theory

  • 基于项目反应理论建模模型与问题的匹配概率
  • 在模型路由和新基准预测上达到当前最优性能
  • 参数可解释,适合需要理解模型能力的场景

近年来大语言模型数量激增,但高效管理和利用这些资源仍是重大挑战。本文探索学习大模型能力的紧凑表示,以支持下游任务如模型路由和新基准上的性能预测。我们将问题建模为:给定模型回答特定问题的正确概率。受心理测量学中的项目反应理论(IRT)启发,将该概率建模为三个关键因素的函数:(i) 模型的多技能能力向量,(ii) 区分不同技能水平模型的问题区分度向量,以及 (iii) 问题难度标量。为联合学习这些参数,我们引入一种混合专家(MoE)网络,耦合模型与问题级别的嵌入。大量实验表明,该方法在模型路由和基准准确率预测上均达到当前最优表现。此外,分析验证了所学参数编码了关于模型能力与问题特征的有意义且可解释的信息。

原文摘要 · Abstract (English)

Recent years have witnessed a surge in the number of large language models (LLMs), yet efficiently managing and utilizing these vast resources remains a significant challenge. In this work, we explore how to learn compact representations of LLM abilities that can facilitate downstream tasks, such as model routing and performance prediction on new benchmarks. We frame this problem as estimating the probability that a given model will correctly answer a specific query. Inspired by the item response theory (IRT) in psychometrics, we model this probability as a function of three key factors: (i) the model's multi-skill ability vector, (2) the query's discrimination vector that separates models of differing skills, and (3) the query's difficulty scalar. To learn these parameters jointly, we introduce a Mixture-of-Experts (MoE) network that couples model- and query-level embeddings. Extensive experiments demonstrate that our approach leads to state-of-the-art performance in both model routing and benchmark accuracy prediction. Moreover, analysis validates that the learned parameters encode meaningful, interpretable information about model capabilities and query characteristics.

大模型评估项目反应理论模型路由能力表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。