arXiv:2507.09090cs.IRcs.CL2025-07被引 2

用大模型开展带检索的辩论,验证其论点与评估能力

DS@GT at Touché: Large Language Models for Retrieval-Augmented Debate

  • 引入检索增强机制,让大模型基于论点库进行结构化辩论
  • 模型回应冗长但评判标准一致,质量与数量表现较好
  • 适合对辩论系统、AI评估机制感兴趣的研究者

大型语言模型(LLMs)展现出强大的对话能力。本文研究其在辩论场景中的双重能力:一是基于论点库进行结构化辩论的能力,二是对辩论中语句进行评估的能力。我们部署了来自三家厂商的六款主流公开模型,参与检索增强型辩论与评估任务。评估采用四个关键指标:质量、数量、表达方式和关联性。实验发现,当提供相关论点时,大模型在辩论中表现良好,但回应普遍冗长,而在评判过程中保持高度一致性。本文配套代码已发布于 https://github.com/dsgt-arc/touche-2025-rad。

原文摘要 · Abstract (English)

Large Language Models (LLMs) demonstrate strong conversational abilities. In this Working Paper, we study them in the context of debating in two ways: their ability to perform in a structured debate along with a dataset of arguments to use and their ability to evaluate utterances throughout the debate. We deploy six leading publicly available models from three providers for the Retrieval-Augmented Debate and Evaluation. The evaluation is performed by measuring four key metrics: Quality, Quantity, Manner, and Relation. Throughout this task, we found that although LLMs perform well in debates when given related arguments, they tend to be verbose in responses yet consistent in evaluation. The accompanying source code for this paper is located at https://github.com/dsgt-arc/touche-2025-rad.

大模型辩论系统检索增强评估机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。