专为政治话语设计的预训练模型,提升分析准确性
RooseBERT: A New Deal For Political Language Modelling
- 在11GB政治辩论语料上预训练,专注政治语言特性
- 在立场识别、论点分类等任务上优于通用模型
- 适合政治分析、舆情监测等研究者使用
政治讨论日益增多,亟需新计算方法自动分析此类内容以促进公民理解。然而,政治语言具有特殊性,且辩论常隐含策略与潜台词,对现有通用预训练语言模型构成挑战。为此,我们提出专用于政治话语的预训练模型RooseBERT,基于11GB英文政治辩论与演讲语料训练。通过在立场检测、情感分析、论点组件识别与分类、论点关系预测、政策分类、命名实体识别(NER)等下游任务上的微调,结果显示其在多数任务中优于通用模型,证明领域特化预训练能显著提升政治话语分析性能。模型已向研究社区开源。
原文摘要 · Abstract (English)
The increasing amount of political debates and politics-related discussions calls for the definition of novel computational methods to automatically analyse such content with the final goal of lightening up political deliberation to citizens. However, the specificity of the political language and the argumentative form of these debates (employing hidden communication strategies and leveraging implicit arguments) make this task very challenging, even for current general-purpose pre-trained Language Models (LMs). To address this, we introduce a novel pre-trained LM for political discourse language called RooseBERT. Pre-training a LM on a specialised domain presents different technical and linguistic challenges, requiring extensive computational resources and large-scale data. RooseBERT has been trained on large political debate and speech corpora (11GB) in English. To evaluate its performances, we fine-tuned it on multiple downstream tasks related to political debate analysis, i.e., stance detection, sentiment analysis, argument component detection and classification, argument relation prediction and classification, policy classification, named entity recognition (NER). Our results show improvements over general-purpose LMs on the majority of these tasks, highlighting how domain-specific pre-training enhances performance in political debate analysis. We release RooseBERT for the research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。