arXiv:2510.15081cs.CLcs.SI2025-10EMNLP被引 1

用大模型自动生成辩论数据,让机器学会识别说服策略。

A Generalizable Rhetorical Strategy Annotation Model Using LLM-based Debate Simulation and Labelling

  • 用大模型模拟辩论生成带标签的合成数据,基于因果、实证、情感、道德四类修辞策略。
  • 模型在多个领域表现优异,跨主题泛化能力强,准确率媲美人工标注。
  • 适合研究政治演讲、广告文案等场景中的说服机制演变。

修辞策略是政治对话、营销宣传和法律论辩中说服力的核心,但现有分析受限于人力标注成本高、不一致且难以扩展。现有数据集多局限于特定话题与策略,制约模型发展。本文提出一种新框架,利用大语言模型(LLM)基于四类修辞类型(因果、实证、情感、道德)自动生成并标注合成辩论数据。在该数据上微调基于Transformer的分类器,并在内部数据集及多个外部语料上验证其性能。结果表明,模型具备高性能与强泛化能力。我们展示了两个应用:(1)引入修辞标签可显著提升说服力预测效果;(2)分析1960–2020年美国总统辩论中修辞策略的时间与党派演变,发现情感型论证使用比例上升,认知型论证下降。

原文摘要 · Abstract (English)

Rhetorical strategies are central to persuasive communication, from political discourse and marketing to legal argumentation. However, analysis of rhetorical strategies has been limited by reliance on human annotation, which is costly, inconsistent, difficult to scale. Their associated datasets are often limited to specific topics and strategies, posing challenges for robust model development. We propose a novel framework that leverages large language models (LLMs) to automatically generate and label synthetic debate data based on a four-part rhetorical typology (causal, empirical, emotional, moral). We fine-tune transformer-based classifiers on this LLM-labeled dataset and validate its performance against human-labeled data on this dataset and on multiple external corpora. Our model achieves high performance and strong generalization across topical domains. We illustrate two applications with the fine-tuned model: (1) the improvement in persuasiveness prediction from incorporating rhetorical strategy labels, and (2) analyzing temporal and partisan shifts in rhetorical strategies in U.S. Presidential debates (1960-2020), revealing increased use of affective over cognitive argument in U.S. Presidential debates.

修辞分析大模型说服力预测政治话语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。