用国会成员推文训练语言模型,实现可预测投票行为的数字孪生。
Toward a digital twin of U.S. Congress
- 基于每位议员任期推文构建动态数据集,训练专属语言模型。
- 生成推文与真实推文几乎无法区分,可准确预测党派投票倾向。
- 适合政策分析、政治预测与资源分配决策者使用。
本文证明,基于语言模型构建的美国国会成员虚拟模型符合数字孪生定义。我们引入一个每日更新的数据集,包含每位国会成员任期内发布的所有推文。研究表明,利用该数据集微调的现代语言模型,能够生成与真实推文高度相似的文本。通过生成推文,可有效预测议员在表决中的行为,并量化其跨党派投票的可能性,为利益相关方提供资源分配支持,或影响实际立法动态。最后讨论了当前分析的局限性及重要拓展方向。
原文摘要 · Abstract (English)
In this paper we provide evidence that a virtual model of U.S. congresspersons based on a collection of language models satisfies the definition of a digital twin. In particular, we introduce and provide high-level descriptions of a daily-updated dataset that contains every Tweet from every U.S. congressperson during their respective terms. We demonstrate that a modern language model equipped with congressperson-specific subsets of this data are capable of producing Tweets that are largely indistinguishable from actual Tweets posted by their physical counterparts. We illustrate how generated Tweets can be used to predict roll-call vote behaviors and to quantify the likelihood of congresspersons crossing party lines, thereby assisting stakeholders in allocating resources and potentially impacting real-world legislative dynamics. We conclude with a discussion of the limitations and important extensions of our analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。