提出评估问答评测集质量的标准化框架,解决评测集自身可信度问题。
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
- 构建元评估框架MEQA,量化评测集质量
- 在网络安全领域测试,发现评测集优劣并对比
- 适合评估LLM评测集的研究者和开发者使用
随着大语言模型(LLMs)的发展,其对社会的潜在影响日益显著,因此严格的模型评估既是技术需求也是社会责任。尽管已有众多评估基准,但对基准本身质量的元评估仍存在关键空白。本文提出MEQA框架,用于对问答类基准进行元评估,实现标准化评分与可比性分析。我们在网络安全领域应用该方法,结合人工与大模型评估者,揭示了现有基准的优势与不足。选择此领域源于AI模型兼具防御工具与安全威胁的双重属性。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) advance, their potential for widespread societal impact grows simultaneously. Hence, rigorous LLM evaluations are both a technical necessity and social imperative. While numerous evaluation benchmarks have been developed, there remains a critical gap in meta-evaluation: effectively assessing benchmarks' quality. We propose MEQA, a framework for the meta-evaluation of question and answer (QA) benchmarks, to provide standardized assessments, quantifiable scores, and enable meaningful intra-benchmark comparisons. We demonstrate this approach on cybersecurity benchmarks, using human and LLM evaluators, highlighting the benchmarks' strengths and weaknesses. We motivate our choice of test domain by AI models' dual nature as powerful defensive tools and security threats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。