arXiv:2411.16077cs.CLcs.MA2024-11被引 1

用智能评估代理提升无参考文本生成质量判断能力

SAGEval: The frontiers of Satisfactory Agent based NLG Evaluation for reference-free open-ended text

  • 引入评阅代理修正大模型评分,无需真实参考文本
  • 在无标注数据下仍可准确评估复杂结构文本生成质量
  • 特别适合评估问卷、表单等开放式生成任务

大型语言模型(LLM)已深度集成于Microsoft365和Google Workspace等应用中,用于文档、邮件、演示文稿等内容的创建与处理,显著提升了工作效率。然而,随着集成复杂度上升,确保输出内容的相关性与适用性变得尤为关键。针对缺乏参考文本或真实标签的自然语言生成(NLG)评估难题,本文提出全新框架SAGEval,利用一个评阅代理对大模型生成的评分进行反馈与修正。实验表明,在无参考/真实标签条件下,该代理能有效纠正大模型评分,显著降低对标注数据的依赖,尤其适用于生成结构化数据如JSON格式表单、多选题、李克特量表、单选题等多样化风格的任务。

原文摘要 · Abstract (English)

Large Language Model (LLM) integrations into applications like Microsoft365 suite and Google Workspace for creating/processing documents, emails, presentations, etc. has led to considerable enhancements in productivity and time savings. But as these integrations become more more complex, it is paramount to ensure that the quality of output from the LLM-integrated applications are relevant and appropriate for use. Identifying the need to develop robust evaluation approaches for natural language generation, wherein references/ground labels doesn't exist or isn't amply available, this paper introduces a novel framework called "SAGEval" which utilizes a critiquing Agent to provide feedback on scores generated by LLM evaluators. We show that the critiquing Agent is able to rectify scores from LLM evaluators, in absence of references/ground-truth labels, thereby reducing the need for labeled data even for complex NLG evaluation scenarios, like the generation of JSON-structured forms/surveys with responses in different styles like multiple choice, likert ratings, single choice questions, etc.

自然语言生成无参考评估大模型评测智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。