arXiv:2604.08797cs.CLcs.AI2026-04被引 1

用多语言故事寓意生成评估大模型的文化适配性

Lessons Without Borders? Evaluating Cultural Alignment of LLMs Using Multilingual Story Moral Generation

  • 构建14个语种-文化组合的故事寓意数据集,以多维度评估模型输出
  • GPT-4o和Gemini生成的寓意与人类高度相似且更受偏好
  • 模型缺乏跨语言差异,集中于少数普适价值,忽视文化多样性

故事是跨文化传播价值观的重要载体,但其解读在语言与文化间存在差异。为此,我们提出多语言故事寓意生成这一新型文化化评估任务。基于涵盖14个语种-文化组合的人类撰写故事寓意数据集,通过语义相似度、人工偏好调查和价值分类,对比模型输出与人类理解。结果表明,GPT-4o与Gemini生成的寓意在语义上与人类响应高度相似,并更受人类评价者青睐。然而,这些模型输出表现出显著降低的跨语言差异,集中在较窄的一组普遍接受的价值观上。这说明当前前沿模型虽能捕捉人类道德理解的共性,却难以复现人类叙事理解中的多样性。本研究将叙事解读转化为可评估任务,为超越静态基准或知识测试,探索语言模型的文化对齐提供了新范式。

原文摘要 · Abstract (English)

Stories are key to transmitting values across cultures, but their interpretation varies across linguistic and cultural contexts. Thus, we introduce multilingual story moral generation as a novel culturally grounded evaluation task. Using a new dataset of human-written story morals collected across 14 language-culture pairs, we compare model outputs with human interpretations via semantic similarity, a human preference survey, and value categorization. We show that frontier models such as GPT-4o and Gemini generate story morals that are semantically similar to human responses and preferred by human evaluators. However, their outputs exhibit markedly less cross-linguistic variation and concentrate on a narrower set of widely shared values. These findings suggest that while contemporary models can approximate central tendencies of human moral interpretation, they struggle to reproduce the diversity that characterizes human narrative understanding. By framing narrative interpretation as an evaluative task, this work introduces a new approach to studying cultural alignment in language models beyond static benchmarks or knowledge-based tests.

大模型评估文化对齐多语言价值观

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。