arXiv:2607.10715cs.CLcs.AI2026-07

构建首个斯拉夫语系修辞说服技术语料库,助力多语言舆论分析。

A Corpus of Persuasion Techniques in Slavic Languages

论文配图:A Corpus of Persuasion Techniques in Slavic Languages
图 1 · 摘自论文原文
  • 基于25类细粒度修辞策略,标注3种斯拉夫语文本中的说服片段。
  • 包含7500个文本片段,覆盖222篇涉及国内外热点议题的文档。
  • 提供机器学习与生成式AI基准,支持跨语言说服分析研究。

说服技巧是广泛应用于各类媒体中影响公众意见的强大修辞手段。本文构建了一个聚焦斯拉夫语系的新型说服技巧语料库,涵盖保加利亚语、波兰语和俄语。语料库在粗粒度文本片段和细粒度句子层面进行了标注,所用技巧源自25种细粒度分类的修辞策略,归入六大类广义说服策略。共包含约7500个文本片段,来自222篇文档,覆盖国内外重大争议话题。我们详细描述了语料库构建流程,提供统计数据,并分析话题与说服技巧间的相关性。采用经典机器学习与生成式AI模型,为文本片段级与句子级说服技巧检测与分类提供基线与基准结果。

原文摘要 · Abstract (English)

Persuasion techniques are powerful rhetorical devices used to sway public opinion in a wide range of media. We present a new corpus of persuasion techniques, focusing on Slavic languages. The corpus contains documents in Bulgarian, Polish, and Russian, annotated with persuasion techniques at the coarse-grained text-span level and fine-grained sentence level. The techniques are drawn from a taxonomy of 25 fine-grained persuasion techniques, grouped under six broad categories of rhetorical persuasion strategies. The corpus contains approximately 7500 text spans from 222 documents that cover topics hotly debated at the national and international levels. We describe the corpus creation process, provide detailed statistics, and examine correlations between topics and persuasion techniques. We use classic ML-based and generative AI-based models to provide baselines and benchmark results for the detection and classification of persuasion techniques at the text-span level and sentence level.

修辞分析语料库多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。