用几何空间统一评估大模型能力,分离真实能力与干扰因素
Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation
- 将九种评估指标转为共享潜空间中的几何度量
- 分解出不稳定性、对齐度、覆盖性三个核心维度
- 适合想设计稳定评测体系的研究者使用
大语言模型评估面临构念效度挑战,碎片化基准和随意指标常将提示敏感性等方法差异与真实能力混淆。现有研究指出,大模型能力与输出可建模为连续的几何流形。本文提出一种广义的多特质多方法(MTMM)框架,系统整合九种评估指标,包括重述不稳定性、漂移分数、奥弗顿宽度和多元性得分,将其解释为共享潜坐标空间中的几何测量。该框架将模型行为分解为三个正交的潜在维度:(1) 不稳定性和敏感性,(2) 位置与对齐度,(3) 覆盖性与表达力。通过系统分离任务无关扰动与真实能力范围,提供理论扎实且领域无关的评测体系设计范式。
原文摘要 · Abstract (English)
The evaluation of Large Language Models (LLMs) faces a critical challenge in construct validity, where fragmented benchmarks and ad hoc metrics frequently conflate method variance, such as prompt sensitivity, with true latent capabilities. Concurrently, emerging research suggests that LLM capabilities and outputs can be modeled as continuous geometric manifolds. In this Systematization of Knowledge (SoK), we bridge these paradigms by proposing a generalized Multi-Trait Multi-Method (MTMM) framework for LLM evaluation. We formalize and unify nine evaluation metrics, including Paraphrase Instability, Drift Score, Overton Width, and Pluralism Score, interpreting them not as isolated scalar values but as geometric measurements within a shared latent coordinate space. This spatial unification factorizes model behavior into three orthogonal latent dimensions: (1) Instability and Sensitivity, (2) Position and Alignment, and (3) Coverage and Expressiveness. By systematically separating task-irrelevant perturbations from true capability spans, the framework provides a theoretically grounded and domain-agnostic taxonomy for robust and empirically stable benchmark design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。