arXiv:2504.00756cs.CL2025-04

用参考数据直接评估大模型知识,效率提升近六成。

RECKON: Large-scale Reference-based Efficient Knowledge Evaluation for Large Language Model

  • 基于参考数据构建结构化评估单元,生成针对性问题。
  • 资源消耗降低56.5%,跨领域准确率超97%。
  • 适合需要高效验证模型知识的开发者和研究者。

随着大语言模型的发展,高效的知识评估变得至关重要。传统方法依赖基准测试,存在资源消耗高、信息损失等问题。本文提出面向大语言模型的大规模参考数据驱动高效知识评估方法RECKON,将非结构化数据组织为可管理单元,并为每个聚类生成针对性问题,显著提升评估准确率与效率。实验表明,RECKON相比传统方法减少56.5%的资源消耗,同时在世界知识、代码、法律及生物医学等多个数据集上实现超过97%的准确率。代码已开源:https://github.com/MikeGu721/reckon。

原文摘要 · Abstract (English)

As large language models (LLMs) advance, efficient knowledge evaluation becomes crucial to verifying their capabilities. Traditional methods, relying on benchmarks, face limitations such as high resource costs and information loss. We propose the Large-scale Reference-based Efficient Knowledge Evaluation for Large Language Model (RECKON), which directly uses reference data to evaluate models. RECKON organizes unstructured data into manageable units and generates targeted questions for each cluster, improving evaluation accuracy and efficiency. Experimental results show that RECKON reduces resource consumption by 56.5% compared to traditional methods while achieving over 97% accuracy across various domains, including world knowledge, code, legal, and biomedical datasets. Code is available at https://github.com/MikeGu721/reckon

知识评估大模型高效评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。