arXiv:2412.05206cs.CLcs.AI2024-12被引 5

用大模型法官评估复杂论证,提升检索增强生成的可信度。

ConQRet: Benchmarking Fine-Grained Evaluation of Retrieval Augmented Argumentation with LLM Judges

  • 设计多维度大模型评委,实现细粒度自动化评估
  • 构建真实网站证据支撑的长篇论证数据集
  • 适合研究者提升论证生成与检索效果评测能力

计算型论证在当今极化环境中日益重要,涉及堕胎禁令、疫苗接种等争议话题。通过检索增强论证(RAArg),大模型可结合真实证据生成有依据的复杂回答。然而,现有评估方法难以应对长文本、高复杂度内容,人工评价成本高且不充分。现有数据集缺乏长篇论证和来自潜在误导性来源的真实证据,无法全面评估检索与论证质量。为此,本文提出使用多个细粒度大模型评委进行自动化评估,相较传统单分制指标及已有众包结果更具可解释性。同时构建新基准ConQRet,包含基于真实网站的长篇人类撰写论证,涵盖检索有效性、论证质量与证据关联性等多维度评测。实验验证了模型评委在旧数据集与ConQRet上的有效性。该方法与基准可加速计算论证研究,并自然扩展至其他检索增强生成任务。

原文摘要 · Abstract (English)

Computational argumentation, which involves generating answers or summaries for controversial topics like abortion bans and vaccination, has become increasingly important in today's polarized environment. Sophisticated LLM capabilities offer the potential to provide nuanced, evidence-based answers to such questions through Retrieval-Augmented Argumentation (RAArg), leveraging real-world evidence for high-quality, grounded arguments. However, evaluating RAArg remains challenging, as human evaluation is costly and difficult for complex, lengthy answers on complicated topics. At the same time, re-using existing argumentation datasets is no longer sufficient, as they lack long, complex arguments and realistic evidence from potentially misleading sources, limiting holistic evaluation of retrieval effectiveness and argument quality. To address these gaps, we investigate automated evaluation methods using multiple fine-grained LLM judges, providing better and more interpretable assessments than traditional single-score metrics and even previously reported human crowdsourcing. To validate the proposed techniques, we introduce ConQRet, a new benchmark featuring long and complex human-authored arguments on debated topics, grounded in real-world websites, allowing an exhaustive evaluation across retrieval effectiveness, argument quality, and groundedness. We validate our LLM Judges on a prior dataset and the new ConQRet benchmark. Our proposed LLM Judges and the ConQRet benchmark can enable rapid progress in computational argumentation and can be naturally extended to other complex retrieval-augmented generation tasks.

论证生成检索增强大模型评估自动化评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。