arXiv:2411.18444cs.CLcs.AI2024-11被引 6

用多模型协作评估会议摘要质量,比人工还准。

Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator

  • 分三步检测错误类型,多智能体讨论优化判断
  • 对错误检测与影响的评分相关性比旧方法高0.25
  • 可自适应定制错误标准,适合数据少的任务

自然语言生成系统生成的会议摘要质量难以自动评估。传统指标如ROUGE和BERTScore与人类判断相关性低,无法捕捉细微错误。近期研究尝试使用大语言模型(LLM),因其具备更好的上下文理解能力,且无需大量人类偏好数据训练即可调整错误定义。然而,现有基于LLM的评估器存在掩盖错误的风险,仅能作为弱代理,仍需人工评估作为金标准,但人工评估成本高且跨研究难以比较。本文提出MESA框架,采用三步评估:逐类错误识别、多智能体讨论优化决策、反馈式自训练以提升错误定义理解与人类判断一致性。以GPT-4o为基底,MESA在错误检测上实现中到高点二列相关系数,误差影响评分达中等斯皮尔曼与肯德尔相关性,平均优于前序方法0.25。该框架可灵活适配定制错误规范,适用于标注数据有限的各类任务。

原文摘要 · Abstract (English)

The quality of meeting summaries generated by natural language generation (NLG) systems is hard to measure automatically. Established metrics such as ROUGE and BERTScore have a relatively low correlation with human judgments and fail to capture nuanced errors. Recent studies suggest using large language models (LLMs), which have the benefit of better context understanding and adaption of error definitions without training on a large number of human preference judgments. However, current LLM-based evaluators risk masking errors and can only serve as a weak proxy, leaving human evaluation the gold standard despite being costly and hard to compare across studies. In this work, we present MESA, an LLM-based framework employing a three-step assessment of individual error types, multi-agent discussion for decision refinement, and feedback-based self-training to refine error definition understanding and alignment with human judgment. We show that MESA's components enable thorough error detection, consistent rating, and adaptability to custom error guidelines. Using GPT-4o as its backbone, MESA achieves mid to high Point-Biserial correlation with human judgment in error detection and mid Spearman and Kendall correlation in reflecting error impact on summary quality, on average 0.25 higher than previous methods. The framework's flexibility in adapting to custom error guidelines makes it suitable for various tasks with limited human-labeled data.

自动评估大模型会议摘要多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。