自动分割并标注政治文本中的立场声明,提升分析效率与泛化能力
Strategies for political-statement segmentation and labelling in unstructured text
- 联合使用分段与分类模型,直接从原始文本提取立场语句
- 在英国议会记录上验证方法,成功追踪三大政党三十年政治轨迹
- 适用于未标注数据,比依赖人工分句的方法更具实用性
对议会演讲和政党宣言的计算分析已成为政治文本研究的重要方向。尽管演讲多采用无监督方法分析,但MARPOR项目已构建了一个包含按条目标注政治立场的大规模宣言语料库。已有研究表明,可通过神经网络预测这些标签,但当前方法依赖预先提供的语句边界,限制了其在新领域的应用。本文提出并测试了一系列统一的分段与标注框架,基于线性链CRF、微调的文本到文本模型,以及上下文学习结合约束解码的组合策略,可直接从原始文本中联合完成语句分割与立场分类。实验表明,该方法在未加工的政治宣言文本上达到有竞争力的准确率,并进一步应用于英国下议院记录,成功追踪四个主要政党的三十年政治演变路径。
原文摘要 · Abstract (English)
Analysis of parliamentary speeches and political-party manifestos has become an integral area of computational study of political texts. While speeches have been overwhelmingly analysed using unsupervised methods, a large corpus of manifestos with by-statement political-stance labels has been created by the participants of the MARPOR project. It has been recently shown that these labels can be predicted by a neural model; however, the current approach relies on provided statement boundaries, limiting out-of-domain applicability. In this work, we propose and test a range of unified split-and-label frameworks -- based on linear-chain CRFs, fine-tuned text-to-text models, and the combination of in-context learning with constrained decoding -- that can be used to jointly segment and classify statements from raw textual data. We show that our approaches achieve competitive accuracy when applied to raw text of political manifestos, and then demonstrate the research potential of our method by applying it to the records of the UK House of Commons and tracing the political trajectories of four major parties in the last three decades.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。