测试大模型在外交决策中的偏见,发现部分模型更倾向军事升级。
Critical Foreign Policy Decisions (CFPD)-Benchmark: Measuring Diplomatic Preferences in Large Language Models
- 设计400个专家构建的国际关系情景,评估7个主流大模型的外交偏好。
- Qwen2、Gemini等模型比Claude、GPT-4o更常推荐军事升级方案。
- 所有模型均对中美俄存在偏好差异,对美英更倾向干预行动。
随着国家安全机构越来越多地将人工智能引入决策与内容生成流程,理解大语言模型(LLMs)的内在偏见至关重要。本研究提出一个新基准,用于评估七种主流基础模型——Llama 3.1 8B Instruct、Llama 3.1 70B Instruct、GPT-4o、Gemini 1.5 Pro-002、Mixtral 8x22B、Claude 3.5 Sonnet 和 Qwen2 72B——在国际关系(IR)背景下的偏见与偏好。我们围绕国际关系核心议题设计了400个专家构建的情景,涵盖军事升级、军事与人道干预、国际体系合作行为及联盟动态四个领域。分析显示,各模型在不同情景下的建议存在显著差异。其中,Qwen2 72B、Gemini 1.5 Pro-002 和 Llama 3.1 8B Instruct 更倾向于推荐激进策略,而 Claude 3.5 Sonnet 与 GPT-4o 则相对保守。所有模型均表现出国家层面的偏见,通常对中国和俄罗斯的激进行动建议较少,而对美国和英国则更支持干预。研究强调,在高风险场景中部署大模型需严格控制,必须进行领域特定评估并针对性微调以对齐机构目标。
原文摘要 · Abstract (English)
As national security institutions increasingly integrate Artificial Intelligence (AI) into decision-making and content generation processes, understanding the inherent biases of large language models (LLMs) is crucial. This study presents a novel benchmark designed to evaluate the biases and preferences of seven prominent foundation models-Llama 3.1 8B Instruct, Llama 3.1 70B Instruct, GPT-4o, Gemini 1.5 Pro-002, Mixtral 8x22B, Claude 3.5 Sonnet, and Qwen2 72B-in the context of international relations (IR). We designed a bias discovery study around core topics in IR using 400-expert crafted scenarios to analyze results from our selected models. These scenarios focused on four topical domains including: military escalation, military and humanitarian intervention, cooperative behavior in the international system, and alliance dynamics. Our analysis reveals noteworthy variation among model recommendations based on scenarios designed for the four tested domains. Particularly, Qwen2 72B, Gemini 1.5 Pro-002 and Llama 3.1 8B Instruct models offered significantly more escalatory recommendations than Claude 3.5 Sonnet and GPT-4o models. All models exhibit some degree of country-specific biases, often recommending less escalatory and interventionist actions for China and Russia compared to the United States and the United Kingdom. These findings highlight the necessity for controlled deployment of LLMs in high-stakes environments, emphasizing the need for domain-specific evaluations and model fine-tuning to align with institutional objectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。