用大模型挖掘说服文本中话语关系的作用,发现六类关系对煽动性表达至关重要。
On the Influence of Discourse Relations in Persuasive Texts
- 用大模型和提示工程为语料标注22类话语关系,构建新型双标签数据集
- 六类话语关系在19种说服技巧中起关键作用,尤其影响夸大、重复等手法
- 成果可助力识别网络宣传与虚假信息,适合传播学与AI交叉研究者参考
本文利用大语言模型(LLMs)与提示工程,探究说服技巧(PTs)与话语关系(DRs)之间的关联。由于缺乏同时标注了PTs与DRs的数据集,研究以包含19种说服技巧的SemEval 2023 Task 3数据集为基础,开发基于LLM的分类器,为每个实例标注22类PDTB 3.0二级话语关系。共评估4个大模型,使用10种提示,生成40个独立的DR分类器。通过不同多数投票策略构建5个银级数据集,规模从1,281到204个实例不等。对这些银级数据集的统计分析显示,六类话语关系(即原因、目的、对比、因果+信念、让步和条件)在说服性文本中起关键作用,尤其与煽动性语言、夸张/最小化、重复及制造怀疑等技巧密切相关。该发现有助于检测网络宣传与虚假信息,并深化对有效沟通机制的理解。
原文摘要 · Abstract (English)
This paper investigates the relationship between Persuasion Techniques (PTs) and Discourse Relations (DRs) by leveraging Large Language Models (LLMs) and prompt engineering. Since no dataset annotated with both PTs and DRs exists, we took the SemEval 2023 Task 3 dataset labelled with 19 PTs as a starting point and developed LLM-based classifiers to label each instance of the dataset with one of the 22 PDTB 3.0 level-2 DRs. In total, four LLMs were evaluated using 10 different prompts, resulting in 40 unique DR classifiers. Ensemble models using different majority-pooling strategies were used to create 5 silver datasets of instances labelled with both persuasion techniques and level-2 PDTB senses. The silver dataset sizes vary from 1,281 instances to 204 instances, depending on the majority pooling technique used. Statistical analysis of these silver datasets shows that six discourse relations (namely Cause, Purpose, Contrast, Cause+Belief, Concession, and Condition) play a crucial role in persuasive texts, especially in the use of Loaded Language, Exaggeration/Minimisation, Repetition and to cast Doubt. This insight can contribute to detecting online propaganda and misinformation, as well as to our general understanding of effective communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。