无需参考答案,用多个独立评估器实现更贴近人类的模型评价。
MILE-RefHumEval: A Reference-Free, Multi-Independent LLM Framework for Human-Aligned Evaluation
- 多个独立提示的评估器协同打分,不依赖人工标注。
- 在对话、摘要等任务中评分与人类高度一致,效率更高。
- 适合需要快速、可靠评估大模型的开发者和研究者。
我们提出MILE-RefHumEval,一种无需参考答案或评估者协调的大型语言模型(LLMs)评估框架。该框架通过一组独立提示的评估器,基于人类对齐的评判标准,支持离散与连续评分。针对最佳候选选择、摘要生成、图像描述及对话等任务,提供灵活、可解释且可扩展的评估方式。实验表明,该方法与人类判断高度一致,优于已有方法,同时降低计算开销,为实际场景中的大模型评估提供了高效、稳健且以人为本的解决方案。
原文摘要 · Abstract (English)
We introduce MILE-RefHumEval, a reference-free framework for evaluating Large Language Models (LLMs) without ground-truth annotations or evaluator coordination. It leverages an ensemble of independently prompted evaluators guided by a human-aligned schema, supporting both discrete and continuous scoring judgement. With task-specific prompts from best candidate selection, summarization and image captioning to dialogue, MILE-RefHumEval provides flexible, interpretable, and scalable assessments. Experiments show it aligns closely with human judgments, outperforms prior methods, and reduces computational overhead, offering an efficient, robust, and human-aligned solution for real-world LLM evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。