arXiv:2511.08066cs.AIcs.CL2025-11被引 7

用文本压缩效率评估大模型推理性能,更准更全面。

Information Capacity: Evaluating the Efficiency of Large Language Models via Text Compression

  • 以文本压缩表现衡量模型效率,融合分词器影响。
  • 56个开源模型测试显示同系列模型压缩能力相近。
  • 适合关注模型资源优化与规模扩展的研究者。

近年来大语言模型(LLM)快速发展,应用广泛,对计算资源的需求持续上升。测试时缩放策略进一步加剧了模型能力与资源消耗之间的矛盾。然而,缺乏一个能准确反映不同分词器、参数量和模型架构下推理效率的严谨度量标准。受压缩与智能之间相关性的启发,我们提出信息容量——一种基于文本压缩性能与计算复杂度比值的模型效率度量。其独特之处在于纳入了分词器效率,该因素影响推理成本但常被忽略。我们评估了56个开源模型的信息容量,发现同一系列中不同规模模型的信息容量保持一致。在五个异构数据集上的实验揭示主流大模型存在显著的语言偏差。实验证明,信息容量可有效预测跨模型规模的性能,并与基准测试得分高度相关。该度量可用于量化推理效率改进,为未来大模型的高效扩展提供洞见。

原文摘要 · Abstract (English)

Recent years have witnessed the rapid advancements of large language models (LLMs) and their expanding applications, leading to soaring demands for computational resources. The widespread adoption of test-time scaling further intensifies the tension between model capability and resource consumption. However, a rigorous metric that accurately reflects an LLM's inference efficiency across diverse tokenizers, parameter counts, and model architectures remains absent. Motivated by the correlation between compression and intelligence, we introduce information capacity, a measure of model efficiency based on text compression performance relative to computational complexity. A distinctive feature of information capacity is its incorporation of tokenizer efficiency, which affects inference costs but is often neglected in LLM evaluations. We assess the information capacity of 56 open-source models and observe a consistent information capacity among different-sized models within a series. Experiments on five heterogeneous datasets reveal strong linguistic biases in mainstream LLMs. Empirical results verify the accuracy of performance prediction across model sizes based on information capacity and show the correlation between information capacity and benchmark scores. This metric can be used to quantify improvements in inference efficiency and provide insights into better scaling performance for future LLM development.

大模型评估推理效率信息容量文本压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。