大模型能识别文化背景却不会据此调整回答,除非被明确提示。
LLMs Infer Cultural Context but Fail to Apply It When Responding

- 通过含文化线索的对话数据,测试模型对文化背景的推理能力。
- 模型能记住文化惯例但常不应用,需显式提示才可适配回应。
- 文化线索越多越能适应,但模型自带西方文化偏好倾向。
近期研究表明,大模型过度代表主流文化(尤其是西方文化),而边缘化其他文化。本文探究这一现象是否影响模型生成文化适配回复的能力,通过评估模型根据用户潜在文化背景使用本地计量单位的情况进行分析。我们构建了包含不同文化线索水平的对话数据集CAPRI。实验表明,当前最先进的大模型虽能推断出文化背景并回忆相关惯例,但通常不会利用这些信息来调整回答,除非被明确提示分步执行任务。进一步评估时间与数量表达的文化解读适应性发现,随着文化线索积累,模型的回答逐渐适应,但其默认倾向并非中立,有时仍偏向模型训练数据来源国(如美国)。整体而言,CAPRI为未来缩小文化知识与文化适应性语言生成之间的差距提供了重要资源。
原文摘要 · Abstract (English)
Recent work has shown that LLMs overrepresent dominant cultures, particularly Western ones, while marginalizing others. We investigate whether this affects models' ability to generate culturally adapted responses by evaluating their use of local measurement units based on the user's perceived cultural background. We introduce Cultural and Pragmatic Response Inference (CAPRI), a dataset of conversations with varying levels of cultural cues. Experiments with state-of-the-art LLMs show that models can infer cultural background and recall relevant conventions, but often fail to utilize the information to adapt their answers to the relevant cultural conventions, unless explicitly prompted to perform the tasks sequentially. We further evaluate adaptation to the interpretation of time and quantity expressions, two subjective language grounding dimensions that are affected by culture. We find that models increasingly adapt their answers as cultural cues accumulate, but their priors are not culture-neutral, sometimes aligning with the model's country of origin. Overall, CAPRI provides a resource for future research aimed at narrowing the gap between cultural knowledge and culturally adaptive language generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。