测试大模型在西非冲突监测中的偏见,发现开源模型易误判合法战斗,适配模型更公正但仍有选择性偏差。
Are LLMs Ready for Conflict Monitoring? Empirical Evidence from West Africa

- 对比开源与领域适配模型在冲突事件分类中的表现,识别出系统性偏见来源。
- 开源模型存在显著虚假非人道化偏差(如Gemma误判18.29%合法战斗),而适配模型接近中立。
- 适配模型仍对国家行为体存在36.5%的过度正当化倾向,需引入人工监督与对抗评估。
随着大语言模型进入冲突监测领域,理解其输出中的系统性偏差对人道主义问责至关重要。我们评估了四种通用开源模型(Gemma 3 4B、Llama 3.2 3B、Mistral 7B、OLMo 2 7B)和两种领域适配模型(AfroConfliBERT、AfroConfliLLAMA)在尼日利亚与喀麦隆冲突事件分类任务上的表现,使用经过多阶段验证的ACLED金标准数据集。结果表明存在方向性分歧:开源模型表现出显著的虚假非人道化偏差——Gemma将18.29%的合法战斗错误归类为针对平民的暴力,且无虚假正当化错误;而两个适配模型则近乎方向中立,正当化偏差差异不显著。然而,领域适配未消除基于行为体的选择性偏差:在尼日利亚,相同战术情境下,国家行为体被正当化比例比非国家行为体高出36.5%。此外,开源模型对地理特异性词汇框架敏感,在喀麦隆中,去合法化短语导致高达66.7%的分类翻转率,在尼日利亚为34.2%;同一扰动在不同语境下影响不同。错误溯源分析显示,模型通过虚构理由掩盖规范性偏差。相比之下,AfroConfliBERT和AfroConfliLLAMA在各类扰动下几乎零翻转率,表现出强鲁棒性。总体而言,当前模型尚不适合在无监督条件下用于冲突监测。我们呼吁开展公平导向微调以降低行为体偏差,强制进行对抗鲁棒性评估以抵御词汇操纵,并根据区域复杂度实施针对性的人机协同监督。
原文摘要 · Abstract (English)
As LLMs enter conflict monitoring, understanding systematic distortions in their outputs is critical for humanitarian accountability. We evaluate four vanilla open-weight models Gemma 3 4B, Llama 3.2 3B, Mistral 7B, and OLMo 2 7B and two domain-adapted models, AfroConfliBERT and AfroConfliLLAMA, on Nigeria and Cameroon conflict-event classification against ACLED, a gold-standard dataset with multi-stage verification. We find a bifurcated divergence in normative directionality. Open-weight models exhibit statistically significant False Illegitimation bias: Gemma misclassifies to 18.29% of legitimate battles as civilian-targeted violence while making zero False Legitimation errors. By contrast, AfroConfliBERT and AfroConfliLLAMA achieve near-directional neutrality, with Legitimization Bias differences indistinguishable from zero. Yet domain adaptation does not eliminate actor-based selection bias. Both adapted models show statistically significant actor bias comparable to vanilla LLMs; in Nigeria, state actors are legitimized 36.5% more often than non-state actors in identical tactical contexts. Open-weight outputs are also fragile to geography-specific lexical framing: delegitimizing phrases produce flip rates up to 66.7% in Cameroon and 34.2% in Nigeria, while perturbations salient in one context may not matter in another. Error trace profiling shows models mask normative bias through unfaithful rationale confabulations. In contrast, AfroConfliBERT and AfroConfliLLAMA are largely robust, with near-zero flip rates across perturbation categories. Overall, current models are not ready for unsupervised deployment in conflict monitoring. We call for fairness-aware fine-tuning to reduce actor-based selection bias, mandatory adversarial robustness evaluation against lexical manipulation, and context-specific human-in-the-loop oversight calibrated to regional difficulty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。