arXiv:2608.25717cs.CLcs.AI2026-08

RAG不能消除地理知识偏差,模型表现仍受自身知识分布影响。

When RAG Fails to Equalize: Geo-bias in Factual Question Answering over Public Companies

论文配图:When RAG Fails to Equalize: Geo-bias in Factual Question Answering over Public Companies
图 1 · 摘自论文原文
  • 在2000家上市公司上测试六种大模型,评估四种上下文下的事实问答能力。
  • 无上下文时模型准确率存在显著地理差异,且完美上下文无法完全弥补差距。
  • 模型越大越准,但知识偏差和错误信息复制问题依然存在,适合关注公平性的研究者看。

检索增强生成(RAG)常被认为能缓解大语言模型(LLMs)的事实性错误,但其是否能均匀补偿模型缺失的知识尚不明确。我们在全球股票指数覆盖的约2000家上市公司上构建了一个可控的事实问答基准,评估六种LLMs在四个原子属性上的表现,涵盖四种情境:无上下文、完美上下文、误导性上下文和干扰性上下文。结果显示,无上下文时模型准确率存在显著地理差异,表明参数化知识分布不均。尽管完美上下文能提升性能,但未能消除这些差距:性能提升与基线准确率正相关,说明检索有效性依赖于内部表征。在误导性上下文中,模型频繁复制错误信息。更大模型虽整体表现更好,但结构性偏差未被消除。结果挑战了RAG作为通用纠错工具的观点,凸显了模型知识、上下文质量与实体表征之间的相互作用。

原文摘要 · Abstract (English)

Retrieval-augmented generation (RAG) is widely assumed to mitigate factual errors in large language models (LLMs), but it remains unclear whether retrieval uniformly compensates for missing knowledge. We study this question in a controlled factual QA setting over public companies, constructing a benchmark of approximately 2,000 firms across global equity indices. We evaluate six LLMs on four atomic attributes under four conditions: no-context, perfect context, misleading context, and distraction context. We find strong geographic disparities in no-context accuracy, indicating uneven parametric knowledge. While perfect context improves performance, it does not eliminate these gaps: gains are correlated with baseline accuracy, suggesting retrieval effectiveness is coupled to internal representations. Under misleading context, models frequently copy incorrect information. Larger models improve overall performance but do not remove these structural effects. These results challenge the view of RAG as a universal corrective and highlight the interaction between model knowledge, context quality, and entity representation.

知识偏差RAG大模型事实问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。