首次用算法分析以色列政治话语中的去合法化言论,发现其30年持续上升。
The Enemy from Within: A Study of Political Delegitimization Discourse in Israeli Political Speech
- 构建1.04万句希伯来语语料库,结合大模型与微调模型识别去合法化言论
- 发现去合法化言论占比17.4%,社交平台比议会发言更常见,右翼和男性政客更频繁使用
- 可自动检测言论强度与情感框架,适合研究民主话语与政治极化
我们首次开展大规模计算研究,分析政治去合法化话语(PDD),即对政治主体规范合法性进行符号性攻击的言论。研究构建并人工标注了包含10,410句的希伯来语语料库,涵盖1993-2023年议会发言、2018-2021年脸书帖子及主流新闻媒体内容,其中1,812例(17.4%)呈现PDD,并对642例额外标注了强度、不文明程度、目标类型与情感框架。提出两阶段分类流程,结合微调编码器与解码大模型,最佳模型DictaLM 2.0在二元PDD检测中达到F₁=0.74,特征分类宏F₁=0.67。应用该模型分析跨平台与纵向数据,发现三十年间PDD显著上升,社交平台高于议会辩论,男性政客使用更多,右翼倾向更强,选举与重大政治事件期间尤为突出。结果表明自动化分析在理解民主话语方面具有可行性与价值。
原文摘要 · Abstract (English)
We present the first large-scale computational study of political delegitimization discourse (PDD), defined as symbolic attacks on the normative validity of political entities. We curate and manually annotate a novel Hebrew-language corpus of 10,410 sentences drawn from Knesset speeches (1993-2023), Facebook posts (2018-2021), and leading news outlets, of which 1,812 instances (17.4\%) exhibit PDD and 642 carry additional annotations for intensity, incivility, target type, and affective framing. We introduce a two-stage classification pipeline combining finetuned encoder models and decoder LLMs. Our best model (DictaLM 2.0) attains an F$_1$ of 0.74 for binary PDD detection and a macro-F$_1$ of 0.67 for classification of delegitimization characteristics. Applying this classifier to longitudinal and cross-platform data, we see a marked rise in PDD over three decades, higher prevalence on social media versus parliamentary debate, greater use by male than female politicians, and stronger tendencies among right-leaning actors - with pronounced spikes during election campaigns and major political events. Our findings demonstrate the feasibility and value of automated PDD analysis for understanding democratic discourse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。