arXiv:2601.07423cs.CL2026-01ACL

构建首个大规模策略性辩论对话数据集,支持多轮互动建模。

SAD: A Large-Scale Strategic Argumentative Dialogue Dataset

  • 基于论证理论标注五类策略,支持多策略共现
  • 包含39万+对话样本,需结合上下文生成适配策略
  • 适合研究交互式推理与说服力生成的学者

论证生成因在人类推理与决策中的核心作用而受到广泛关注。然而,现有论证语料库大多聚焦于非交互、单轮场景,或从给定话题生成论点,或反驳已有论点。实际上,论证常以多轮对话形式展开,发言者需维护立场并运用多样化的论证策略提升说服力。为支持更深入的论证对话建模,我们提出首个大规模战略型论证对话数据集SAD,包含392,822个样本。基于论证理论,我们对每个话语标注五类策略类型,允许多策略共现。与以往数据集不同,SAD要求模型生成需结合对话历史、指定立场及目标策略的上下文相关论点。我们进一步对多种预训练生成模型在SAD上进行基准测试,并深入分析了论证中的策略使用模式。

原文摘要 · Abstract (English)

Argumentation generation has attracted substantial research interest due to its central role in human reasoning and decision-making. However, most existing argumentative corpora focus on non-interactive, single-turn settings, either generating arguments from a given topic or refuting an existing argument. In practice, however, argumentation is often realized as multi-turn dialogue, where speakers defend their stances and employ diverse argumentative strategies to strengthen persuasiveness. To support deeper modeling of argumentation dialogue, we present the first large-scale \textbf{S}trategic \textbf{A}rgumentative \textbf{D}ialogue dataset, SAD, consisting of 392,822 examples. Grounded in argumentation theories, we annotate each utterance with five strategy types, allowing multiple strategies per utterance. Unlike prior datasets, SAD requires models to generate contextually appropriate arguments conditioned on the dialogue history, a specified stance on the topic, and targeted argumentation strategies. We further benchmark a range of pretrained generative models on SAD and present in-depth analysis of strategy usage patterns in argumentation.

对话生成论证生成大规模数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。