用多智能体自动评估大模型,解决评测数据过时与人工依赖问题。
MACEval: A Multi-Agent Continual Evaluation Network for Large Models
- 构建多智能体网络,实现动态自主评估
- 在23个大模型上验证有效,显著降低评估开销
- 适合关注持续评测与自动化评估的研究者
近年来已有数百个针对大模型的评测基准发布,但多数为封闭式评测,易因数据污染导致过拟合。此外,当前评测集规模与范围持续扩大,指标瞬变,且高度依赖人工维护,难以及时更新。本文提出 MACEval,一种用于大模型动态评估的多智能体持续评估网络,定义新指标以量化性能长期变化。MACEval 采用交互式自主评估模式,通过角色分配、过程内数据生成和级联智能体路由实现评估流程自动化。在23个大模型上的大量实验表明,MACEval 有效减轻评估负担,显著降低运行开销。我们希望该框架能推动大模型评测的未来方向。项目主页:https://github.com/zijianchen98/MACEval。
原文摘要 · Abstract (English)
Hundreds of benchmarks dedicated to evaluating large models have been presented over the past few years. However, most of them remain closed-ended and are prone to overfitting due to the potential data contamination. Moreover, the increasing scale and scope of current benchmarks with transient metrics, as well as the heavily human-dependent curation procedure, pose significant challenges for timely maintenance and adaptation. In this paper, we introduce MACEval, a Multi-Agent Continual Evaluation network for dynamic evaluation of large models, and define new metrics to quantify performance longitudinally. MACEval employs an interactive and autonomous evaluation mode, utilizing role assignment, in-process data generation, and evaluation routing through a cascaded agent network. Extensive experiments on 23 large models demonstrate the effectiveness of MACEval, which also lightens the evaluation process and reduces a considerable amount of overhead. We hope that MACEval can broaden future directions of large model evaluation. Project page: https://github.com/zijianchen98/MACEval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。