用熵和矩阵核范数融合评估大模型,兼顾精度与效率
Combining Entropy and Matrix Nuclear Norm for Enhanced Evaluation of Language Models
- 结合协方差矩阵熵与矩阵核范数,从隐藏状态中提取不确定性与冗余信息
- 在多个LLM上验证,复合评分能更稳定反映模型性能差异
- 可调节权重适配不同评估目标,适合模型对比与诊断
随着大语言模型(LLMs)的持续发展,精确高效的评估指标需求日益迫切。传统方法虽具信息量,但常受限于计算开销与可解释性。本文提出一种新型混合评估方法,融合协方差矩阵熵与矩阵核范数(MNN)。首先对LLM的隐藏状态进行归一化,再基于其计算协方差矩阵与MNN;进一步计算协方差矩阵的熵以捕捉输出中的不确定性与冗余性。通过将两者组合为综合评分,构建兼顾准确性与计算效率的评估框架。该方法支持灵活调节熵与MNN的权重,适用于不同评估目标。在多种LLM上的实验验证了方法的鲁棒性与有效性,为模型性能分析提供了更深入视角。本工作推动了LLM评估技术的发展,并为未来模型评测创新开辟路径。
原文摘要 · Abstract (English)
As large language models (LLMs) continue to advance, the need for precise and efficient evaluation metrics becomes more pressing. Traditional approaches, while informative, often face limitations in computational demands and interpretability. In this paper, we introduce a novel hybrid evaluation method that integrates two established techniques: entropy derived from covariance matrices and the Matrix Nuclear Norm (MNN). Our method begins by normalizing hidden states from LLMs, then computes the covariance matrix and MNN from these representations. We further calculate the entropy of the covariance matrix to capture uncertainty and redundancy in the model's outputs. By combining these metrics into a composite score, we offer a comprehensive evaluation framework that balances accuracy with computational efficiency. Additionally, our approach allows for flexibility in adjusting the weightings between entropy and MNN, tailoring the evaluation for different objectives. Through a series of experiments on various LLMs, we demonstrate the robustness and efficacy of our method, offering deeper insights into model performance. This work contributes to the ongoing development of LLM evaluation and opens avenues for future innovations in model assessment techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。