arXiv:2502.18874cs.CLcs.AI2025-02ACL被引 1

让大模型自动评估更准更稳,支持多维度分析与代码验证。

Learning to Align Multi-Faceted Evaluation: A Unified and Robust Framework

  • 自适应生成评价标准,融合文本与代码双重分析
  • 在多种指令下表现更稳定,优于现有微调评估器
  • 适合需要高可靠评估的AI系统开发与测试场景

大型语言模型(LLMs)正被广泛用于各类场景的自动化评估。以往研究尝试微调开源LLM以复现如GPT-4等闭源模型的评估解释与判断,但大多局限于预设通用标准下的文本分析,对未见指令适应性差,且在评估定量与结构约束时表现不稳定。为此,我们提出新型评估框架ARJudge,可自适应生成评价标准,并融合文本与代码驱动分析来评估LLM输出。ARJudge包含两个组件:用于生成多维度评估分析的微调分析器(Analyzer),以及无需微调的精炼器(Refiner),用于整合并优化所有分析结果以得出最终判断。我们构建了一个综合分析语料库(Composite Analysis Corpus),涵盖评价标准生成、文本分析与代码分析任务,用于训练分析器。实验表明,ARJudge在有效性和鲁棒性上均优于现有微调评估器,同时验证了多维度评估与代码驱动分析对提升评估能力的关键作用。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are being used more and more extensively for automated evaluation in various scenarios. Previous studies have attempted to fine-tune open-source LLMs to replicate the evaluation explanations and judgments of powerful proprietary models, such as GPT-4. However, these methods are largely limited to text-based analyses under predefined general criteria, resulting in reduced adaptability for unseen instructions and demonstrating instability in evaluating adherence to quantitative and structural constraints. To address these limitations, we propose a novel evaluation framework, ARJudge, that adaptively formulates evaluation criteria and synthesizes both text-based and code-driven analyses to evaluate LLM responses. ARJudge consists of two components: a fine-tuned Analyzer that generates multi-faceted evaluation analyses and a tuning-free Refiner that combines and refines all analyses to make the final judgment. We construct a Composite Analysis Corpus that integrates tasks for evaluation criteria generation alongside text-based and code-driven analysis generation to train the Analyzer. Our results demonstrate that ARJudge outperforms existing fine-tuned evaluators in effectiveness and robustness. Furthermore, it demonstrates the importance of multi-faceted evaluation and code-driven analyses in enhancing evaluation capabilities.

大模型评估多维度分析代码驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。