arXiv:2410.10724cs.CL2024-10被引 7

让大模型主动发现任务需求,自动给出符合人类判断的生成评价。

Large Language Models Are Active Critics in NLG Evaluation

  • 大模型自主推断任务类型与评价标准,无需人工预设规则。
  • 在多任务场景下,评分与人类判断一致性显著提升。
  • 适合需要灵活适应不同评价场景的生成系统评估。

传统基于大语言模型(LLM)的自然语言生成(NLG)评估依赖预定义的任务定义和评价标准,使大模型成为遵循开发者指令的“被动批评者”。然而,人类评估者常使用隐含标准,其期望会因具体用户需求而异。因此,现有方法在缺乏大量提示定制的情况下难以适应多样化场景。为此,我们提出 Active-Critic,一种新型基于大模型的评估框架,将大模型转变为能利用少量示例数据自适应多种NLG任务的“主动批评者”。该框架包含两个阶段:(1) 自主推断目标NLG任务及相关评价标准;(2) 动态优化提示以生成与人类判断对齐的评分及详细理由。实验表明,Active-Critic 能生成细致、上下文感知的评价标准,在多个任务上实现了与人类判断更高的对齐度。

原文摘要 · Abstract (English)

The conventional paradigm of using large language models (LLMs) for natural language generation (NLG) evaluation relies on pre-defined task definitions and evaluation criteria, positioning LLMs as "passive critics" that strictly follow developer-provided guidelines. However, human evaluators often apply implicit criteria, and their expectations in practice can vary widely based on specific end-user needs. Consequently, these rigid evaluation methods struggle to adapt to diverse scenarios without extensive prompt customization. To address this, we introduce Active-Critic, a novel LLM-based evaluator that transforms LLMs into "active critics'' capable of adapting to diverse NLG tasks using limited example data. Active-Critic consists of two stages: (1) self-inferring the target NLG task and relevant evaluation criteria, and (2) dynamically optimizing prompts to produce human-aligned scores along with detailed justifications. Our experiments show that Active-Critic can generate nuanced, context-aware evaluation criteria, enabling it to achieve superior alignment with human judgments across multiple tasks.

NLG评估大模型主动评价

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。