用语言复杂度评估做零样本代理,快速判断大模型能力。
Language Complexity Measurement as a Noisy Zero-Shot Proxy for Evaluating LLM Performance
- 用LIX和依赖距离衡量模型语言分析能力
- ChatGPT-o1-mini在两项任务中表现最稳定
- 语言复杂度与综合能力高度相关,适合快速评估
大型语言模型在自然语言生成方面取得显著进展,但在需要精确计算和结构分析的任务中仍面临挑战。本文通过计算瑞典中学与大学作文的LIX可读性指标和平均依赖距离(ADD),评估了先进LLM在语言复杂度测量任务中的表现,并与已知真实值对比。结果表明,所有模型均具备一定能力,其中ChatGPT-o1-mini表现最为一致,在LIX计算和依存句法分析中准确率最高。此外,模型在LIX计算上的准确率与MMLU基准测试的总体表现存在显著负相关(r = -0.875, p = 0.026, N=6)。这些发现表明,语言复杂度测量可作为评估大模型通用能力的噪声零样本代理,无需依赖大量基准数据集即可实现高效模型评估。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have made significant strides in natural language generation but often face challenges in tasks requiring precise calculations and structural analysis. This paper investigates the performance of state-of-the-art LLMs on language complexity measurement tasks, through the computation of the LIX readability metric and Average Dependency Distance (ADD). Using Swedish high school and university-level essays, we evaluate the models' abilities to compute LIX scores and perform dependency parsing, comparing their results to established ground truths. Our findings reveal that while all models demonstrate some capacity for these tasks, ChatGPT-o1-mini performs most consistently, achieving the highest accuracy in both LIX computation and dependency parsing. Additionally, we observe a strong significant correlation -0.875 p 0.026 (N=6) between the models' accuracy in computing LIX and their overall performance on the Massive Multitask Language Understanding (MMLU) benchmark. These results suggest that language complexity measurement abilities can serve as a noisy zero-shot proxies for assessing the general capabilities of LLMs, providing a practical method for model evaluation without the need for extensive benchmarking datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。