简单把政党信息加在演讲前,比复杂方法更有效。
Language Models Learn Metadata: Political Stance Detection Case Study
- 将政党、政策等元数据直接拼接在演讲文本前
- 仅用政党信息的基线模型超越当前最优方法
- 适合关注政治立场分析与元数据融合的研究者
立场检测是自然语言处理中的关键任务,广泛应用于社会科学研究,如分析在线讨论和评估政治宣传。本文研究如何最优地将元数据融入政治立场检测任务。我们发现,以往将元数据与语言数据结合的方法未能充分挖掘元数据信息;简单的基线模型——仅使用政党成员信息——已超越当前最先进的方法。进一步实验表明,将元数据(如政党、政策)前置到政治演讲文本前效果最佳,优于所有基线,说明复杂的元数据整合系统可能并未最优学习该任务。
原文摘要 · Abstract (English)
Stance detection is a crucial NLP task with numerous applications in social science, from analyzing online discussions to assessing political campaigns. This paper investigates the optimal way to incorporate metadata into a political stance detection task. We demonstrate that previous methods combining metadata with language-based data for political stance detection have not fully utilized the metadata information; our simple baseline, using only party membership information, surpasses the current state-of-the-art. We then show that prepending metadata (e.g., party and policy) to political speeches performs best, outperforming all baselines, indicating that complex metadata inclusion systems may not learn the task optimally.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。