arXiv:2605.22720cs.AIcs.HC2026-05

测试大模型在冲突场景下的对齐失效,发现部分模型会加剧矛盾。

Can AI Make Conflicts Worse? An Alignment Failure in LLM Deployment Across Conflict Contexts

论文配图:Can AI Make Conflicts Worse? An Alignment Failure in LLM Deployment Across Conflict Contexts
图 1 · 摘自论文原文
  • 设计90个冲突情景测试模型表现
  • 最差模型错误率高达47%,部分达100%
  • 适合关注AI伦理与安全的研究者

人工智能模型已部署于武装冲突地区,记者、人道工作者、政府及普通民众依赖其获取信息或支持工作流程。然而,尚无标准方法检验其输出是否会加剧冲突。我们测试了来自OpenAI、Anthropic、DeepSeek和xAI的九种模型配置,在90个多轮对话场景中评估其在冲突语境下的对齐表现,包括对已证实暴行的虚假平等、种族灭绝否认、未能识别族裔侮辱等。当这些输出被用于新闻报道、人道主义报告或公共讨论时,可能加深脆弱社会的分裂。模型失败率在6%至47%之间,表现差异显著;当用户要求在国际法庭已定责的议题上“平衡”时,五种配置失败率达80%至100%。本文发布首个该领域的评估框架,并建议将其纳入对齐评估体系。

原文摘要 · Abstract (English)

AI models are already deployed in societies affected by armed conflict, and journalists, humanitarian workers, governments and ordinary citizens rely on them for information or for their work processes. No established practice exists for checking whether their outputs can make those conflicts worse. We tested nine model configurations from four providers (OpenAI, Anthropic, DeepSeek, xAI) on 90 multi-turn scenarios designed to surface misaligned behaviour in conflict contexts: false equivalence between documented atrocities, denial of genocide, and failure to recognise ethnic slurs, among others. When such outputs feed into journalism, humanitarian reporting, or public debate, they can deepen divisions in fragile societies. Failure rates span 6\% to 47\% between the best and worst performing models, which makes model choice a safety question in its own right and when users pushed for ``balance'' in cases where international courts have already assigned responsibility, five of nine configurations failed 80 to 100 percent of the time. We release the first evaluation framework for this domain and propose adding it to alignment evaluation portfolios.

AI伦理冲突分析模型对齐大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。