SPEED框架通过专家反馈提升大模型评估的公平性与可解释性
Integrated Framework for LLM Evaluation with Answer Generation
- 引入专用专家模型进行多维度描述性分析
- 在多个数据集上表现稳定,且资源消耗更低
- 适合关注评估公正性与透明度的研究者
可靠的大语言模型评估对实际应用至关重要。传统基于基准的评估方法依赖固定参考答案,难以捕捉生成内容的重要定性特征。为此,我们提出集成评估框架SPEED(自修正描述性评估与专家驱动诊断),利用专业功能专家对模型输出进行全方位、描述性分析。不同于传统方法,SPEED在幻觉检测、毒性评估和词汇-上下文适当性等多个维度主动融入专家反馈。实验表明,SPEED在多样领域和数据集上均实现稳健一致的评估性能。此外,采用相对紧凑的专家模型,其资源效率优于大规模评估器。结果表明,SPEED显著提升了大模型评估的公平性与可解释性,为现有评估方法提供了有力替代方案。
原文摘要 · Abstract (English)
Reliable evaluation of large language models is essential to ensure their applicability in practical scenarios. Traditional benchmark-based evaluation methods often rely on fixed reference answers, limiting their ability to capture important qualitative aspects of generated responses. To address these shortcomings, we propose an integrated evaluation framework called \textit{self-refining descriptive evaluation with expert-driven diagnostics}, SPEED, which utilizes specialized functional experts to perform comprehensive, descriptive analyses of model outputs. Unlike conventional approaches, SPEED actively incorporates expert feedback across multiple dimensions, including hallucination detection, toxicity assessment, and lexical-contextual appropriateness. Experimental results demonstrate that SPEED achieves robust and consistent evaluation performance across diverse domains and datasets. Additionally, by employing relatively compact expert models, SPEED demonstrates superior resource efficiency compared to larger-scale evaluators. These findings illustrate that SPEED significantly enhances fairness and interpretability in LLM evaluations, offering a promising alternative to existing evaluation methodologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。