arXiv:2510.22823cs.CLcs.AI2025-10被引 5

对比商业与开源大模型在人道主义多语言监测中的稳定性与偏见。

Cross-Lingual Stability and Bias in Instruction-Tuned Language Models for Humanitarian NLP

  • 用多语言推理评估六模型,量化跨语言可靠性
  • 对齐模型在低资源语言上表现稳定,开源模型易受语言影响
  • 为预算有限组织提供可靠模型选型建议

人道主义组织面临抉择:投入昂贵商业API或使用免费开源模型进行多语言人权监测。尽管商业系统更可靠,但开源模型缺乏实证验证,尤其在冲突地区常见的低资源语言中。本文首次系统比较商业与开源大语言模型(LLMs)在七种语言中的人权侵害检测能力,通过78,000次多语言推理评估六模型——四款指令对齐模型(Claude-Sonnet-4、DeepSeek-V3、Gemini-Flash-2.0、GPT-4.1-mini)和两款开源模型(LLaMA-3-8B、Mistral-7B),采用标准分类指标及新提出的跨语言可靠性度量:校准偏差(CD)、决策偏倚(B)、语言鲁棒性得分(LRS)和语言稳定性得分(LSS)。结果表明,对齐程度决定稳定性:对齐模型在语系差异大且低资源语言(如林加拉语、缅甸语)中保持近似不变的准确率与平衡校准,而开源模型表现出显著的提示语言敏感性和校准漂移。研究证明多语言对齐可实现语言无关推理,并为资源受限组织在预算与可靠性间权衡提供实用指导。

原文摘要 · Abstract (English)

Humanitarian organizations face a critical choice: invest in costly commercial APIs or rely on free open-weight models for multilingual human rights monitoring. While commercial systems offer reliability, open-weight alternatives lack empirical validation -- especially for low-resource languages common in conflict zones. This paper presents the first systematic comparison of commercial and open-weight large language models (LLMs) for human-rights-violation detection across seven languages, quantifying the cost-reliability trade-off facing resource-constrained organizations. Across 78,000 multilingual inferences, we evaluate six models -- four instruction-aligned (Claude-Sonnet-4, DeepSeek-V3, Gemini-Flash-2.0, GPT-4.1-mini) and two open-weight (LLaMA-3-8B, Mistral-7B) -- using both standard classification metrics and new measures of cross-lingual reliability: Calibration Deviation (CD), Decision Bias (B), Language Robustness Score (LRS), and Language Stability Score (LSS). Results show that alignment, not scale, determines stability: aligned models maintain near-invariant accuracy and balanced calibration across typologically distant and low-resource languages (e.g., Lingala, Burmese), while open-weight models exhibit significant prompt-language sensitivity and calibration drift. These findings demonstrate that multilingual alignment enables language-agnostic reasoning and provide practical guidance for humanitarian organizations balancing budget constraints with reliability in multilingual deployment.

多语言人权监测模型稳定性开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。