用AI识别中俄维基百科的说服性语言差异,揭示文化视角区别。
Uncovering Differences in Persuasive Language in Russian versus English Wikipedia
- 将说服检测转为大模型自生成的高层问题,提升识别准确性。
- 在8.8万篇双语维基文章中发现俄版更强调乌克兰议题,英版偏中东。
- 方法可跨语言应用,适合研究文化差异与信息传播的学者。
我们研究英语与俄语维基百科文章中说服性语言的差异,以揭示不同文化对各类主题的独特视角。提出一种基于大语言模型(LLM)的系统,通过让模型自动生成高层问题(HLQs)来识别说服性内容,而非直接判断说服性。首先由LLM生成大量候选问题,再经人工标签筛选出小规模精准集合。采用两阶段‘识别-提取’提示策略,在包含88,000篇文章的大规模双语数据集上定位说服实例。量化每篇文章的说服程度,并对比成对文章的差异。结果显示,按说服度排序的文章与文化直觉一致:俄语维基突出乌克兰议题,英语维基聚焦中东。将主题归类后发现政治事件比其他类别更具说服性。此外,HLQ在两种语言中表现相当。该方法支持大规模跨语言、跨文化理解,代码、提示和数据均已开源。
原文摘要 · Abstract (English)
We study how differences in persuasive language across Wikipedia articles, written in either English and Russian, can uncover each culture's distinct perspective on different subjects. We develop a large language model (LLM) powered system to identify instances of persuasive language in multilingual texts. Instead of directly prompting LLMs to detect persuasion, which is subjective and difficult, we propose to reframe the task to instead ask high-level questions (HLQs) which capture different persuasive aspects. Importantly, these HLQs are authored by LLMs themselves. LLMs over-generate a large set of HLQs, which are subsequently filtered to a small set aligned with human labels for the original task. We then apply our approach to a large-scale, bilingual dataset of Wikipedia articles (88K total), using a two-stage identify-then-extract prompting strategy to find instances of persuasion. We quantify the amount of persuasion per article, and explore the differences in persuasion through several experiments on the paired articles. Notably, we generate rankings of articles by persuasion in both languages. These rankings match our intuitions on the culturally-salient subjects; Russian Wikipedia highlights subjects on Ukraine, while English Wikipedia highlights the Middle East. Grouping subjects into larger topics, we find politically-related events contain more persuasion than others. We further demonstrate that HLQs obtain similar performance when posed in either English or Russian. Our methodology enables cross-lingual, cross-cultural understanding at scale, and we release our code, prompts, and data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。