arXiv:2505.17793cs.CL2025-05被引 3

发现语言模型压缩假象,用几何分析提升智能评估准确性

Compression Hacking: A Supplementary Perspective on Informatics Properties of Language Models from Geometric Distortion

  • 通过几何畸变分析修正压缩率度量方法
  • 新指标与模型综合能力相关性超0.9,显著优于旧方法
  • 适合关注模型内在结构与评估可靠性的研究者

近期「压缩即智能」为语言模型提供了新的信息学度量视角,强调高度结构化的表征体现模型智能水平。但从几何角度看,高度压缩的语言模型其词向量空间易退化为高度各向异性状态,削弱指令理解能力并直接影响性能。我们发现压缩与各向异性同步现象本质上是模型表征中的「压缩欺骗」:噪声主导方向通过牺牲空间均匀性制造出高压缩的假象。基于此,我们提出三种融合几何畸变分析的改进压缩度量,并集成至自评估流水线。新度量与模型综合能力的斯皮尔曼相关系数超过0.9,显著优于原始压缩度量及其他基于内部结构的度量。结果表明,引入表征几何畸变可实质性增强语言模型的信息学解释力。

原文摘要 · Abstract (English)

Recently, the concept of ``compression as intelligence'' has provided a novel informatics metric perspective for language models (LMs), emphasizing that highly structured representations signify the intelligence level of LMs. However, from a geometric standpoint, the word representation space of highly compressed LMs tends to degenerate into a highly anisotropic state, which hinders the LM's ability to comprehend instructions and directly impacts its performance. We found this compression-anisotropy synchronicity is essentially the ``Compression Hacking'' in LM representations, where noise-dominated directions tend to create the illusion of high compression rates by sacrificing spatial uniformity. Based on this, we propose three refined compression metrics by incorporating geometric distortion analysis and integrate them into a self-evaluation pipeline. The refined metrics exhibit strong alignment with the LM's comprehensive capabilities, achieving Spearman correlation coefficients above 0.9, significantly outperforming both the original compression and other internal structure-based metrics. This confirms that compression hacking substantially enhances the informatics interpretation of LMs by incorporating geometric distortion of representations.

语言模型压缩评估几何分析模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。