arXiv:2502.14476cs.CL2025-02

构建首个比较问答评估基准,量化大模型生成质量。

Argument-Based Comparative Question Answering Evaluation Benchmark

  • 设计15项人工与大模型标注的评估标准
  • Llama-3 70B在摘要评价中表现最佳,GPT-4答问能力最强
  • 公开数据、代码与结果,支持复现与对比

本文针对自动比较问答评估中的挑战,提出一个评估框架以衡量比较问答摘要的质量。我们基于人工标注及6个大语言模型和两个比较问答数据集的标注,制定了15项评估标准。在多种设置下对多个大语言模型与人工标注进行测试,验证了评估的一致性。结果表明,Llama-3 70B Instruct在摘要评价中表现最佳,而GPT-4在回答比较问题上领先。所有数据、代码与评估结果均已公开。

原文摘要 · Abstract (English)

In this paper, we aim to solve the problems standing in the way of automatic comparative question answering. To this end, we propose an evaluation framework to assess the quality of comparative question answering summaries. We formulate 15 criteria for assessing comparative answers created using manual annotation and annotation from 6 large language models and two comparative question asnwering datasets. We perform our tests using several LLMs and manual annotation under different settings and demonstrate the constituency of both evaluations. Our results demonstrate that the Llama-3 70B Instruct model demonstrates the best results for summary evaluation, while GPT-4 is the best for answering comparative questions. All used data, code, and evaluation results are publicly available\footnote{\url{https://anonymous.4open.science/r/cqa-evaluation-benchmark-4561/README.md}}.

问答评估大模型评测比较问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。