测试几何指标对大模型生成质量的评估能力,发现其主要受文本长度影响。
Geometric Metrics and LLMs: What They Measure and When They Work
- 通过六种模型和八类任务,系统检验八种几何指标的可靠性。
- 部分指标(如Schatten Norm)实质反映输出长度,控制后区分能力消失。
- 结合文本统计特征可提升生成模型识别准确率至78%,适合故障检测场景。
我们对大语言模型评估中的几何度量进行了系统的压力测试。基于内部表示的秩相关几何特性作为无参考的质量信号展现出潜力,但其可靠条件尚不明确。我们在六种测试模型(0.5-8B)和八类生成任务上,评估了八种常用度量:内在维度估计器、谱范数及相关量,分离出真正的几何信号、文本长度效应以及标准文本统计已捕捉的内容。三个发现浮现:第一,某些度量(特别是Schatten Norm和MOM)主要反映输出长度,一旦控制长度,其判别力即崩溃;第二,几何度量在文本统计之外提供了微弱但真实的信息:与文本统计结合,分类器在6类生成器识别上达到78%准确率,高于文本统计单独使用的69%;第三,内在维度与词汇多样性(RTTR)之间的关联性仅中等。我们提出针对性使用建议,并指出故障检测是最具前景的近期应用场景。
原文摘要 · Abstract (English)
We present a systematic stress-test of geometric metrics for LLM evaluation. Rank-based geometric properties of internal representations have shown promise as reference-free quality signals, but the conditions under which they are reliable remain unclear. We evaluate eight commonly-used metrics: intrinsic-dimensionality estimators, spectral norms, and related quantities across six tester models (0.5-8B) and eight generators on contrasting tasks, separating genuine geometric signal from text-length effects and from what standard text statistics already capture. Three findings emerge. First, some metrics (notably Schatten Norm and MOM) mainly reflect output length, and their apparent discriminative power collapses once length is controlled. Second, geometric metrics add modest but real information beyond text statistics: combined with them, a classifier reaches 78% accuracy on 6-way generator identification versus 69% for text statistics alone. Third, rather than tracking a general notion of text quality, the metrics demonstrate only moderate association between the intrinsic-dimensionality and lexical diversity (RTTR). We give use-case-specific recommendations and identify failure detection as the most promising near-term application.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。