arXiv:2502.13497cs.CLcs.AI2025-02ACL被引 10

用网络搜索提升大模型文化认知,但可能强化刻板印象。

Towards Geo-Culturally Grounded LLM Generations

  • 用网络搜索检索增强大模型,提升对文化事实的认知能力。
  • 搜索增强使多选题得分显著提高,但人类评估中文化熟悉度未改善。
  • 提醒我们区分文化知识与文化素养,警惕模型刻板化风险。

生成式大语言模型在全球多元文化认知方面存在差距。本文研究检索增强生成与网络搜索增强技术对模型文化意识的影响。通过对比标准大模型、基于自建知识库的检索增强模型(KB grounding)以及基于网络搜索的检索增强模型(search grounding),在多个文化认知评测基准上进行测试。结果发现,搜索增强显著提升了模型在测试文化规范、文物与制度等命题知识的多选题表现;而知识库增强受限于知识覆盖不足和检索器效果不佳。然而,搜索增强也增加了模型产生刻板判断的风险,在具备足够统计效力的人类评估中,未能提升评委对模型文化熟悉度的认可。这揭示了命题性文化知识与开放性文化素养之间的差异,对评估大模型的文化意识具有重要启示。

原文摘要 · Abstract (English)

Generative large language models (LLMs) have demonstrated gaps in diverse cultural awareness across the globe. We investigate the effect of retrieval augmented generation and search-grounding techniques on LLMs' ability to display familiarity with various national cultures. Specifically, we compare the performance of standard LLMs, LLMs augmented with retrievals from a bespoke knowledge base (i.e., KB grounding), and LLMs augmented with retrievals from a web search (i.e., search grounding) on multiple cultural awareness benchmarks. We find that search grounding significantly improves the LLM performance on multiple-choice benchmarks that test propositional knowledge (e.g., cultural norms, artifacts, and institutions), while KB grounding's effectiveness is limited by inadequate knowledge base coverage and a suboptimal retriever. However, search grounding also increases the risk of stereotypical judgments by language models and fails to improve evaluators' judgments of cultural familiarity in a human evaluation with adequate statistical power. These results highlight the distinction between propositional cultural knowledge and open-ended cultural fluency when it comes to evaluating LLMs' cultural awareness.

大模型文化认知检索增强刻板印象

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。