首个面向大模型论辩能力的标准化评测基准,覆盖46项任务
ArgBench: Benchmarking LLMs on Computational Argumentation Tasks

- 构建统一33数据集的论辩任务评测基准
- 验证少样本、推理步骤等对性能的关键影响
- 适合评估大模型在辩论、批判性思维中的应用能力
论辩能力是大型语言模型(LLMs)的重要工具。该能力在自我反思、协同辩论以获取多样答案以及反驳仇恨言论等场景中至关重要。本文首次构建了面向基于大模型的计算论辩任务的标准化评测基准,整合了来自前期研究的33个数据集。基于该基准,我们评估了五大类大模型家族在46项计算论辩任务上的泛化能力,涵盖论点挖掘、观点评估、论据质量判断、论据推理和论点生成。我们还系统分析了少量示例、推理步骤、模型规模及训练策略对模型性能的影响。
原文摘要 · Abstract (English)
Argumentation skills are an essential toolkit for large language models (LLMs). These skills are crucial in various use cases, including self-reflection, debating collaboratively for diverse answers, and countering hate speech. In this paper, we create the first benchmark for a standardized evaluation of LLM-based approaches to computational argumentation, encompassing 33 datasets from previous work in unified form. Using the benchmark, we evaluate the generalizability of five LLM families across 46 computational argumentation tasks that cover mining arguments, assessing perspectives, assessing argument quality, reasoning about arguments, and generating arguments. On the benchmark, we conduct an extensive systematic analysis of the contribution of few-shot examples, reasoning steps, model size, and training skills to the performance of LLMs on the computational argumentation tasks in the benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。