arXiv:2609.08459cs.CL2026-09

用风格分析法识别政治文本中的隐藏作者,助力政策研究与传播分析。

Detecting Authorship in Political Texts with Inductive Stylometry

论文配图:Detecting Authorship in Political Texts with Inductive Stylometry
图 1 · 摘自论文原文
  • 结合字符三元组与降维技术,通过风格差异定位潜在作者
  • 在英匈双语法律文本中识别出几乎不重叠的写作风格指纹
  • 适合研究立法过程、政治传播及演讲稿背后的真实作者

政治文本极少由名义发言人独自撰写。推文、演讲、报告和官方声明常由幕僚起草、修改或协调,但政治学对此类隐性作者留下的风格痕迹关注不足。本文提出并严苛测试一种归纳式风格分析方法,结合字符3-gram特征与UMAP降维,辅以Burrows' Delta。该方法应用于六个语料库,涵盖从推文到长篇文档、书面与口头形式、英语与匈牙利语。结果表明,该方法在两种语言的正式法律文本中成功识别出近乎不重叠的分析师风格指纹;能将政客推文合理分组并揭示额外信息;可区分脚本化与即兴发言。但在脚本化语料中无法分辨具体撰稿人。频率型风格分析是强大工具,其有效性取决于作者信号强度与机构编辑程度,适用于立法研究、政治传播与政策分析。

原文摘要 · Abstract (English)

Political texts are rarely authored by the nominal speaker alone. Tweets, speeches, reports, and official statements are drafted, edited, or harmonized by staff, yet political science has paid limited attention to the stylistic traces these hidden authors leave behind. This paper develops and stress-tests an inductive stylometric approach for recovering latent authorship structure in political communication, combining character 3-gram features with UMAP dimensionality reduction, and Burrows' Delta. We apply the approach to six corpora that vary in length (from tweets to long documents), in mode (written and oral), and in language (English and Hungarian). The approach recovers near-disjoint analyst fingerprints in formal legal prose in both languages, sorts a politician's tweets into validated subsets while uncovering additional insights, and distinguishes scripted from improvised speech. It fails, however, to resolve individual speechwriters within scripted corpora. Frequency-based stylometry is thus a powerful tool that, depending on authorial signal strength and institutional editing, can uncover authorship traces relevant to legislative studies, political communication, and policy research.

风格分析政治传播作者识别自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。