用情境感知的多维测评模型预测大模型在新问题上的表现
LLM Evaluation on Unseen Questions: Contextual Multidimensional IRT Model

- 结合问题上下文与多维项目反应理论建模模型能力
- 在同场景下预测准确率优于无模型基线,多维结构更细致刻画能力差异
- 跨场景泛化仍存挑战,适合关注模型评估效率与可解释性的研究者
大语言模型(LLMs)的评估日益需要在未见问题或任务上预测模型表现,而无需大量新标注。该问题具有挑战性,因题目难度、场景及底层能力需求差异显著。简单回顾平均值可能混淆模型能力与题目特性。本文提出一种基于模型的评估框架,将多维项目反应理论(Multidimensional IRT)与问题上下文结合,以预测模型在未见问题上的表现。该框架通过潜在能力特征表示模型,利用问题内容推断题目特性,实现信息超越已观测题目的迁移。实证结果表明,在同场景评估中,引入问题嵌入能提升预测性能,多维潜结构比单维方法更丰富地描述能力变化。同时,研究揭示一个重要局限:泛化能力并不必然带来跨场景下的可靠预测。这些发现表明,情境感知的心理测量建模是高效且可解释的LLM评估的有前景方向,但也凸显了跨场景泛化这一核心开放挑战。
原文摘要 · Abstract (English)
Evaluation of large language models (LLMs) increasingly requires predicting how a model will perform on new questions or tasks before collecting large amounts of new annotations. This problem is challenging because question difficulty, scenario, and underlying capability demands can vary substantially. Simple retrospective averages may confound model ability with item characteristics. In this paper, we study a model-based evaluation framework that combines multidimensional item response theory model with question contexts to predict LLM performance on unseen questions. The framework represents LLMs through latent capability profiles while using question content to inform item characteristics, allowing information to transfer beyond previously observed items. Empirically, we find that for within-scenario evaluation, incorporating question embeddings improves prediction relative to model-free baselines, and that multidimensional latent structure provides a richer description of capability variation than unidimensional alternatives. At the same time, our results reveal an important limitation that the generalizability does not necessarily translate into reliable prediction under cross-scenario shift. These findings suggest that context-aware psychometric modeling is a promising direction for efficient and interpretable LLM evaluation, while also highlighting cross-scenario generalization as a central open challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。