用新系统发现中文大模型能力分布不均,能高效诊断其专长短板。
ReLE: A Scalable System and Structured Benchmark for Diagnosing Capability Anisotropy in Chinese LLMs
- 构建可扩展评估系统,通过符号化混合评分消除推理任务误判。
- 计算成本降70%仍保持96%排名相关性,覆盖超20万样本。
- 揭示模型排名易受权重影响,适合关注模型真实专长的研究者。
大型语言模型在中文理解上进展迅速,但评估仍受限于基准饱和与高昂算力。本文提出ReLE(鲁棒高效动态评估)系统,用于诊断模型在不同领域间的能力异质性。我们评估了304个模型(189个商用、115个开源),覆盖由207,843个样本构成的领域×能力正交矩阵。提出两项方法创新:(1) 符号基础混合评分机制,消除推理任务中基于嵌入的虚假阳性;(2) 基于奈曼分配与噪声修正的动态方差感知调度器,相比全量评估降低70%算力,同时保持ρ=0.96的排名相关性。分析显示,聚合排名对权重高度敏感:在ReLE中模型秩稳定性幅度(RSA)达11.4,远高于传统基准的~5.0,证实现代模型高度专业化而非全面领先。ReLE并非替代静态基准,而是作为动态监测模型演进的高频诊断工具。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have achieved rapid progress in Chinese language understanding, yet accurately evaluating their capabilities remains challenged by benchmark saturation and prohibitive computational costs. While static leaderboards provide snapshot rankings, they often mask the structural trade-offs between capabilities. In this work, we present ReLE (Robust Efficient Live Evaluation), a scalable system designed to diagnose Capability Anisotropy, the non-uniformity of model performance across domains. Using ReLE, we evaluate 304 models (189 commercial, 115 open-source) across a Domain $\times$ Capability orthogonal matrix comprising 207,843 samples. We introduce two methodological contributions to address current evaluation pitfalls: (1) A Symbolic-Grounded Hybrid Scoring Mechanism that eliminates embedding-based false positives in reasoning tasks; (2) A Dynamic Variance-Aware Scheduler based on Neyman allocation with noise correction, which reduces compute costs by 70\% compared to full-pass evaluations while maintaining a ranking correlation of $ρ=0.96$. Our analysis reveals that aggregate rankings are highly sensitive to weighting schemes: models exhibit a Rank Stability Amplitude (RSA) of 11.4 in ReLE versus $\sim$5.0 in traditional benchmarks, confirming that modern models are highly specialized rather than generally superior. We position ReLE not as a replacement for comprehensive static benchmarks, but as a high-frequency diagnostic monitor for the evolving model landscape.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。