arXiv:2502.14122cs.CLcs.CY2025-02AAAI被引 11

首个面向联合国决策的LLM评测基准,测试模型在政治研判与外交模拟中的表现。

Benchmarking LLMs for Political Science: A United Nations Perspective

  • 构建涵盖1994-2024年联合国安理会记录的全新数据集
  • 提出四项政治科学任务,覆盖起草、投票、讨论全流程
  • 适用于政策研究、国际关系与AI治理方向的研究者

大型语言模型(LLMs)在自然语言处理领域取得显著进展,但其在高风险政治决策中的潜力仍待探索。本文聚焦联合国(UN)决策过程,引入一个包含1994至2024年联合国安理会(UNSC)公开记录的新数据集,涵盖决议草案、投票记录和外交发言。基于此,我们提出首个综合性评估基准——联合国基准(UNBench),用于衡量LLMs在四个相互关联的政治科学任务上的表现:共同提案国判断、代表投票模拟、决议通过预测和代表声明生成。这些任务覆盖了联合国决策过程的三个阶段——起草、投票与讨论,旨在评估模型对政治动态的理解与模拟能力。实验分析揭示了在该领域应用LLMs的潜力与挑战,为人工智能与政治科学的交叉研究提供了新路径。相关代码与数据可在https://github.com/yueqingliang1/UNBench 获取。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved significant advances in natural language processing, yet their potential for high-stake political decision-making remains largely unexplored. This paper addresses the gap by focusing on the application of LLMs to the United Nations (UN) decision-making process, where the stakes are particularly high and political decisions can have far-reaching consequences. We introduce a novel dataset comprising publicly available UN Security Council (UNSC) records from 1994 to 2024, including draft resolutions, voting records, and diplomatic speeches. Using this dataset, we propose the United Nations Benchmark (UNBench), the first comprehensive benchmark designed to evaluate LLMs across four interconnected political science tasks: co-penholder judgment, representative voting simulation, draft adoption prediction, and representative statement generation. These tasks span the three stages of the UN decision-making process--drafting, voting, and discussing--and aim to assess LLMs' ability to understand and simulate political dynamics. Our experimental analysis demonstrates the potential and challenges of applying LLMs in this domain, providing insights into their strengths and limitations in political science. This work contributes to the growing intersection of AI and political science, opening new avenues for research and practical applications in global governance. The UNBench Repository can be accessed at: https://github.com/yueqingliang1/UNBench.

政治决策联合国语言模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。