arXiv:2608.08143cs.IRcs.CL2026-08中稿 · publication in the…

用检索增强的LLM生成辩论回应并评估其质量

DS@GT ARC at Touché: Large Language Models for Retrieval-Augmented Debate

论文配图:DS@GT ARC at Touché: Large Language Models for Retrieval-Augmented Debate
图 1 · 摘自论文原文
  • 六款主流大模型通过检索增强提示生成辩论回复
  • 模型间评价共识高,但与官方评分差异显著,尤其在质量维度
  • 适合关注大模型辩论生成与评价可靠性的研究者

我们扩展了DS@GT ARC的工作笔记,提交至Touché 2025检索增强辩论任务。该任务包含两个子任务:生成模拟辩论中的下一句发言,以及根据格莱斯量、质、相关、方式准则评估辩论回复。本次提交采用来自三个供应商的六款领先大模型,通过检索增强提示管道实现。我们总结了工作论文结果,并探讨多模型评估者一致性是否可作为官方评估性能的可靠代理。分析表明,前沿大模型在生成回复方面表现强劲,同一模型家族内评估者一致性高。然而,这种共识并未可靠反映官方评估目标,尤其在质量准则上差距最大。附带代码见https://github.com/dsgt-arc/touche-2025-rad和https://github.com/dsgt-arc/touche-2025-rad-analysis。

原文摘要 · Abstract (English)

We extend the DS@GT ARC working-note submission to the Touché 2025 Retrieval-Augmented Debate task. The task has two subtasks: generating the next utterance in a simulated debate, and evaluating debate responses according to the Gricean maxims of Quantity, Quality, Relation, and Manner. The DS@GT ARC submission consisted of six leading LLMs from three providers through a retrieval-augmented prompting pipeline. We summarize the results from the working paper and explore whether multi-LLM evaluator agreement is a reliable proxy for official evaluation performance. The analysis shows that frontier LLM systems are strong response generators, and as evaluators they agree strongly within model families. However this consensus does not reliably track the official evaluation target, with the largest gap on the Quality maxim. The accompanying source code for this paper is located at https://github.com/dsgt-arc/touche-2025-rad and https://github.com/dsgt-arc/touche-2025-rad-analysis.

大模型辩论生成评估一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。