提出新框架与数据集,系统解析宣传的技巧、情感诉求与意图。
PropaInsight: Toward Deeper Understanding of Propaganda in Terms of Techniques, Appeals, and Intent
- 基于社会科学研究,将宣传分解为技巧、诉求和意图三要素。
- 合成数据集使模型在技巧识别上提升203.4%,诉求分析提升66.2%。
- 适用于数据稀缺或跨领域场景,适合反虚假信息研究者使用。
宣传在塑造公众舆论和传播虚假信息中起关键作用。现有研究多聚焦于识别宣传技巧,却难以捕捉其深层动机与实际影响。为此,我们提出 propainsight,一个基于基础社会科学的理论框架,系统地将宣传拆解为技巧、情感唤起诉求和潜在意图。该框架提供更细致的跨情境分析能力。同时,我们构建了 propagaze,一个结合人工标注与高质量合成数据的新数据集,其生成流程经过精心设计。实验表明,主流大模型在宣传分析上表现不佳,但使用 propagaze 训练后性能显著提升:微调后的 Llama-7B-Chat 在技巧识别上的文本跨度交并比(IoU)提升 203.4%,在诉求分析中的 BertScore 提升 66.2%,远超 1-shot GPT-4-Turbo。此外,propagaze 在数据稀疏和跨领域场景中有效补充有限的人工标注数据,展现出全面且泛化的宣传分析潜力。
原文摘要 · Abstract (English)
Propaganda plays a critical role in shaping public opinion and fueling disinformation. While existing research primarily focuses on identifying propaganda techniques, it lacks the ability to capture the broader motives and the impacts of such content. To address these challenges, we introduce propainsight, a conceptual framework grounded in foundational social science research, which systematically dissects propaganda into techniques, arousal appeals, and underlying intent. propainsight offers a more granular understanding of how propaganda operates across different contexts. Additionally, we present propagaze, a novel dataset that combines human-annotated data with high-quality synthetic data generated through a meticulously designed pipeline. Our experiments show that off-the-shelf LLMs struggle with propaganda analysis, but training with propagaze significantly improves performance. Fine-tuned Llama-7B-Chat achieves 203.4% higher text span IoU in technique identification and 66.2% higher BertScore in appeal analysis compared to 1-shot GPT-4-Turbo. Moreover, propagaze complements limited human-annotated data in data-sparse and cross-domain scenarios, showing its potential for comprehensive and generalizable propaganda analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。