arXiv:2606.24610cs.CL2026-06被引 1

测试大模型跨文化叙事时,发现它们保义不保形,故事结构差异明显。

Same Lesson, Different Story: Cross-Lingual Reconstruction of Cultural Narratives in Large Language Models

论文配图:Same Lesson, Different Story: Cross-Lingual Reconstruction of Cultural Narratives in Large Language Models
图 1 · 摘自论文原文
  • 用414则跨语言谚语作提示,生成1.3万条故事进行对比
  • 模型保留谚语核心含义,但主角和情节结构系统性改变
  • 不同模型在多语言下趋同,说明共享语义抽象能力

当多种文化传递相同道德教训时,评估文化根基语境变得复杂。该研究提出一种多语言叙事评估框架,整合了15种语言的414则谚语,并使用4个大语言模型生成13,000条叙事。通过语义等价的谚语作为文化根基提示,分析模型是否在跨语言中保持意义、跨语言提示如何影响叙事实现,以及不同模型家族是否收敛于相似解释。结果表明,跨语言提示基本保持谚语级语义意义,但系统性地改变了代理关系、社会定位和叙事结构。此外,在单语和跨语言设置下均观察到强模型间一致性,表明多语言大模型虽架构与语言各异,却依赖共享语义抽象。这揭示了现有评估需更全面考虑文化根基。仅依赖语义相似性可能高估文化保真度,忽视表达中的文化差异。

原文摘要 · Abstract (English)

The evaluation of cultural grounding context becomes complex when multiple cultures convey the same moral lesson. This challenge is particularly relevant to large language models (LLMs), which produce narratives across a wide range of languages and cultural contexts. However, it remains uncertain whether these models preserve culturally grounded meaning when equivalent moral lessons are conveyed through distinct cultural forms. This study introduces a multilingual evaluation narrative framework that integrates a cross-linguistic collection of 414 proverbs spanning 15 languages and uses four LLMs to generate 13k narratives. By employing semantically equivalent proverbs as culturally grounded prompts, the analysis assesses whether models preserve meaning across languages, how cross-lingual conditioning influences narrative realization, and whether different model families converge on similar interpretations. Results indicate that cross-lingual prompting largely preserves proverb-level semantic meaning while systematically redistributing agency, social positioning, and narrative structure. Additionally, strong inter-model convergence is observed in both monolingual and cross-lingual settings, suggesting that multilingual LLMs rely on shared semantic abstractions despite architectural and linguistic differences. These findings shed light on the need for more comprehensive evaluations of cultural grounding. Relying exclusively on semantic similarity in multilingual narrative assessments may overestimate cultural preservation by neglecting culturally meaningful variations in narrative expression.

文化建模跨语言叙事生成大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。