用少量文化样本就能提升大模型跨文化推理能力。
Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World
- 仅需12个文化样本,即可让多语言模型在其他阿拉伯国家提升10%表现。
- 非本地文化示例(如印尼、美国)效果不输本地样本。
- 为低资源文化场景的模型适配提供高效新路径。
大型语言模型常体现西方中心偏见,限制其在多元文化背景下的应用。尽管已有研究关注文化对齐,但利用一文化对齐来提升另一文化表现的跨文化迁移潜力仍待探索。本文聚焦阿拉伯世界,基于覆盖13个阿拉伯国家的文化基准常识推理数据集,评估了轻量级对齐方法(如上下文学习和基于示范的强化,DITTO),并与监督微调、直接偏好优化等基线对比。结果表明,仅需12个来自某国的文化特定示例,即可使多语言模型在其他国家平均提升10%的性能;此外,来自印尼和美国的跨文化示范在多项选择题推理任务中表现可媲美甚至超越本地示范,证明文化常识具有跨区域可迁移性。研究揭示了高效跨文化对齐的可能性,为低资源文化环境中的模型适应提供了可行方案。
原文摘要 · Abstract (English)
Large language models (LLMs) often reflect Western-centric biases, limiting their effectiveness in diverse cultural contexts. Although some work has explored cultural alignment, the potential for cross-cultural transfer, using alignment in one culture to improve performance in others, remains underexplored. This paper investigates cross-cultural transfer of commonsense reasoning in the Arab world, where linguistic and historical similarities coexist with local cultural differences. Using a culturally grounded commonsense reasoning dataset covering 13 Arab countries, we evaluate lightweight alignment methods such as in-context learning and demonstration-based reinforcement (DITTO), alongside baselines like supervised fine-tuning and direct preference optimization. Our results show that merely 12 culture-specific examples from one country can improve performance in others by 10\% on average, within multilingual models. In addition, we demonstrate that out-of-culture demonstrations from Indonesia and US contexts can match or surpass in-culture alignment for MCQ reasoning, highlighting cultural commonsense transferability beyond the Arab world. These findings demonstrate that efficient cross-cultural alignment is possible and offer a promising approach to adapt LLMs to low-resource cultural settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。