让大模型主动发现任务需求,自动给出符合人类判断的生成评价。
Large Language Models Are Active Critics in NLG Evaluation
- 大模型自主推断任务类型与评价标准,无需人工预设规则。
- 在多任务场景下,评分与人类判断一致性显著提升。
- 适合需要灵活适应不同评价场景的生成系统评估。
传统基于大语言模型(LLM)的自然语言生成(NLG)评估依赖预定义的任务定义和评价标准,使大模型成为遵循开发者指令的“被动批评者”。然而,人类评估者常使用隐含标准,其期望会因具体用户需求而异。因此,现有方法在缺乏大量提示定制的情况下难以适应多样化场景。为此,我们提出 Active-Critic,一种新型基于大模型的评估框架,将大模型转变为能利用少量示例数据自适应多种NLG任务的“主动批评者”。该框架包含两个阶段:(1) 自主推断目标NLG任务及相关评价标准;(2) 动态优化提示以生成与人类判断对齐的评分及详细理由。实验表明,Active-Critic 能生成细致、上下文感知的评价标准,在多个任务上实现了与人类判断更高的对齐度。
原文摘要 · Abstract (English)
The conventional paradigm of using large language models (LLMs) for natural language generation (NLG) evaluation relies on pre-defined task definitions and evaluation criteria, positioning LLMs as "passive critics" that strictly follow developer-provided guidelines. However, human evaluators often apply implicit criteria, and their expectations in practice can vary widely based on specific end-user needs. Consequently, these rigid evaluation methods struggle to adapt to diverse scenarios without extensive prompt customization. To address this, we introduce Active-Critic, a novel LLM-based evaluator that transforms LLMs into "active critics'' capable of adapting to diverse NLG tasks using limited example data. Active-Critic consists of two stages: (1) self-inferring the target NLG task and relevant evaluation criteria, and (2) dynamically optimizing prompts to produce human-aligned scores along with detailed justifications. Our experiments show that Active-Critic can generate nuanced, context-aware evaluation criteria, enabling it to achieve superior alignment with human judgments across multiple tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。