arXiv:2503.04809cs.CLcs.AI2025-03被引 1

提出三种方法提升大模型无参考评估效果

PanguIR Technical Report for NTCIR-18 AEOLLM Task

  • 用多模型协作模拟人类评分,覆盖不同任务子项
  • 通过反馈迭代优化提示词,提升评估准确性
  • 结合语义检索与实例选择,优化上下文学习效果

随着大语言模型在学术界和产业界受到广泛关注,其能力的有效评估变得日益关键且具有挑战性。现有评估方法可分为人工评估和自动评估两类:人工评估虽全面但成本高;自动评估虽可扩展,却受限于以参考答案为主的评价标准。为应对这一挑战,NTCIR-18引入了AEOLLM(大模型自动评估)任务,旨在推动无参考评估方法的发展。本文针对该任务,提出三项核心改进策略:1)多模型协作,利用多个大模型在各类子任务中逼近人类评分;2)提示词自动优化,基于训练样本的评估反馈,迭代优化初始任务提示;3)上下文学习(ICL)优化,基于多任务评估反馈,训练专用的上下文示例检索模型,联合语义相关性检索模型,精准筛选最优上下文学习示例。在最终数据集上的实验表明,所提方法在AEOLLM任务中表现优异。

原文摘要 · Abstract (English)

As large language models (LLMs) gain widespread attention in both academia and industry, it becomes increasingly critical and challenging to effectively evaluate their capabilities. Existing evaluation methods can be broadly categorized into two types: manual evaluation and automatic evaluation. Manual evaluation, while comprehensive, is often costly and resource-intensive. Conversely, automatic evaluation offers greater scalability but is constrained by the limitations of its evaluation criteria (dominated by reference-based answers). To address these challenges, NTCIR-18 introduced the AEOLLM (Automatic Evaluation of LLMs) task, aiming to encourage reference-free evaluation methods that can overcome the limitations of existing approaches. In this paper, to enhance the evaluation performance of the AEOLLM task, we propose three key methods to improve the reference-free evaluation: 1) Multi-model Collaboration: Leveraging multiple LLMs to approximate human ratings across various subtasks; 2) Prompt Auto-optimization: Utilizing LLMs to iteratively refine the initial task prompts based on evaluation feedback from training samples; and 3) In-context Learning (ICL) Optimization: Based on the multi-task evaluation feedback, we train a specialized in-context example retrieval model, combined with a semantic relevance retrieval model, to jointly identify the most effective in-context learning examples. Experiments conducted on the final dataset demonstrate that our approach achieves superior performance on the AEOLLM task.

大模型评估无参考评估提示优化多模型协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。