arXiv:2505.16408cs.CL2025-05EMNLP被引 9

用百科和叙事数据提升大模型的文化差异表达能力

From Surveys to Narratives: Rethinking Cultural Value Adaptation in LLMs

  • 融合维基百科与情景叙事数据,补充问卷式文化数据
  • 仅用问卷数据会弱化文化差异,影响事实知识
  • 叙事数据能更好保留文化独特性,适合任务导向应用

大语言模型在适配文化价值观时面临偏见和训练数据不足的挑战。以往研究主要依赖世界价值观调查(WVS)数据进行对齐,但其是否有效捕捉文化细微差别仍不明确,且可能影响下游任务表现。本文系统评估基于WVS的训练方法,发现仅使用问卷数据会导致文化规范同质化,并干扰事实知识。为此,我们引入来自维基百科和NormAd的百科知识与情景化文化叙事作为补充。尽管这些叙事对不同下游任务的影响存在差异,但总体上显著提升了文化区分度。研究揭示了将文化价值观与具体任务行为对齐的内在复杂性。代码已开源:https://github.com/faridlazuarda/from-surveys-to-narratives。

原文摘要 · Abstract (English)

Adapting cultural values in Large Language Models (LLMs) presents significant challenges, particularly due to biases and limited training data. Prior work primarily aligns LLMs with different cultural values using World Values Survey (WVS) data. However, it remains unclear whether this approach effectively captures cultural nuances or produces distinct cultural representations for various downstream tasks. In this paper, we systematically investigate WVS-based training for cultural value adaptation and find that relying solely on survey data can homogenize cultural norms and interfere with factual knowledge. To investigate these issues, we augment WVS with encyclopedic and scenario-based cultural narratives from Wikipedia and NormAd. While these narratives may have variable effects on downstream tasks, they consistently improve cultural distinctiveness than survey data alone. Our work highlights the inherent complexity of aligning cultural values with the goal of guiding task-specific behavior. We release our code at https://github.com/faridlazuarda/from-surveys-to-narratives.

文化对齐大模型叙事数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。