arXiv:2512.11110cs.CLcs.AI2025-12被引 1

构建多语言事实推理基准,发现提示语言会影响模型偏见。

FIBER: A Multilingual Evaluation Resource for Factual Inference Bias

  • 设计多语言事实推理任务,覆盖单/多实体场景。
  • 31%主题存在推理偏差(得分>0.5),土耳其提示偏差更高。
  • 大模型表现更优,多实体问题难度显著高于单实体。

大型语言模型广泛应用,但其事实可靠性与偏见引发关注。现有评测多聚焦单实体、单语数据。为此,我们提出FIBER,一个涵盖英语、意大利语和土耳其语的多语言基准,包含句子补全、问答和对象计数任务,评估单实体与多实体场景下的事实知识。实验表明,提示语言会影响模型在实体选择上的推理结果,尤其当实体关联对应国家时。31%的主题中,事实推理偏差得分超过0.5。不同主题下,土耳其提示导致的偏差高于意大利语,出现在83%的主题中,显示语言依赖性。多实体问题比单实体问题更难处理,模型性能随语言和规模差异显著:英语表现最优,土耳其与意大利语得分明显更低;大模型(如Llama-3.1-8B、Qwen-2.5-7B)优于3B-4B小模型。

原文摘要 · Abstract (English)

Large language models are widely used across domains, yet there are concerns about their factual reliability and biases. Factual knowledge probing offers a systematic means to evaluate these aspects. Most existing benchmarks focus on single-entity facts and monolingual data. We therefore present FIBER, a multilingual benchmark for evaluating factual knowledge in single- and multi-entity settings. The dataset includes sentence completion, question-answering, and object-count prediction tasks in English, Italian, and Turkish. Using FIBER, we examine whether the prompt language induces inference bias in entity selection and how large language models perform on multi-entity versus single-entity questions. The results indicate that the language of the prompt can influence the model's generated output, particularly for entities associated with the country corresponding to that language. However, this effect varies across different topics such that 31% of the topics exhibit factual inference bias score greater than 0.5. Moreover, the level of bias differs across languages such that Turkish prompts show higher bias compared to Italian in 83% of the topics, suggesting a language-dependent pattern. Our findings also show that models face greater difficulty when handling multi-entity questions than the single-entity questions. Model performance differs across both languages and model sizes. The highest mean average precision is achieved in English, while Turkish and Italian lead to noticeably lower scores. Larger models, including Llama-3.1-8B and Qwen-2.5-7B, show consistently better performance than smaller 3B-4B models.

多语言推理偏见评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。