arXiv:2509.25138cs.CL2025-09Conference of the …被引 2

发现多语言模型在事实核查中存在语言与检索双重偏见

Investigating Language and Retrieval Bias in Multilingual Previously Fact-Checked Claim Detection

  • 用20种语言测试6个开源多语言模型,评估跨语言表现差异
  • 热门观点被频繁检索,导致性能虚高,冷门观点被忽略
  • 提示策略和模型规模影响偏见程度,需优化以提升公平性

多语言大模型在跨语言事实核查中展现强大能力,但常表现出语言偏见,对英语等高资源语言表现优于低资源语言。本文提出并分析一种新概念——检索偏见:信息检索系统倾向于偏向某些信息,导致检索结果失衡。研究聚焦于已核查声明检测(PFCD),在20种语言上采用全多语言提示策略,利用AMC-16K数据集评估6个开源多语言LLM。通过将任务提示翻译至各语言,揭示了单语与跨语言性能差异,并识别出模型家族、规模与提示策略的关键影响趋势。研究还使用多语言嵌入模型分析检索频率,发现部分声明在不同帖子中被过度检索,导致热门声明的检索性能被夸大,而较少出现的声明则被低估。结果表明,多语言模型行为中仍存在显著偏见,为提升多语言事实核查的公平性提供改进方向。

原文摘要 · Abstract (English)

Multilingual Large Language Models (LLMs) offer powerful capabilities for cross-lingual fact-checking. However, these models often exhibit language bias, performing disproportionately better on high-resource languages such as English than on low-resource counterparts. We also present and inspect a novel concept - retrieval bias, when information retrieval systems tend to favor certain information over others, leaving the retrieval process skewed. In this paper, we study language and retrieval bias in the context of Previously Fact-Checked Claim Detection (PFCD). We evaluate six open-source multilingual LLMs across 20 languages using a fully multilingual prompting strategy, leveraging the AMC-16K dataset. By translating task prompts into each language, we uncover disparities in monolingual and cross-lingual performance and identify key trends based on model family, size, and prompting strategy. Our findings highlight persistent bias in LLM behavior and offer recommendations for improving equity in multilingual fact-checking. To investigate retrieval bias, we employed multilingual embedding models and look into the frequency of retrieved claims. Our analysis reveals that certain claims are retrieved disproportionately across different posts, leading to inflated retrieval performance for popular claims while under-representing less common ones.

多语言事实核查偏见分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。