arXiv:2502.18583cs.CL2025-02EMNLP被引 2

测试大模型对后苏联地区菜肴的文化认知,发现其常误判菜系归属。

What are Foundation Models Cooking in the Post-Soviet World?

  • 构建俄语、乌克兰语双语多模态数据集BORSch,包含1147道俄罗斯菜和823道乌克兰菜。
  • 主流模型在文本和图文问答中错误地将菜肴归因于提问语言对应的国家。
  • 揭示语言混用与训练数据偏差是导致文化认知偏差的关键原因。

后苏联地区的文化复杂且受历史动荡持续影响。本研究通过构建涵盖1147道俄罗斯菜和823道乌克兰菜的多模态数据集BORSch,考察基础模型对该区域饮食文化的知识水平。实验表明,领先模型在纯文本与图文问答任务中均难以准确识别菜肴来源国,反而倾向于将菜肴归因于提问所用语言对应的国家。通过对预训练数据的分析,我们发现这一现象源于菜品与国家名称的误导性共现,以及俄语与乌克兰语之间的代码混用等语言现象。为突破问答评估的局限性,我们进一步测试模型生成菜肴视觉描述的能力。结果显示该任务与问答表现相关性弱,暗示仅依赖问答可能无法全面评估模型的文化理解力。为促进后续研究,我们将BORSch公开发布于https://github.com/alavrouk/BORSch。

原文摘要 · Abstract (English)

The culture of the Post-Soviet states is complex, shaped by a turbulent history that continues to influence current events. In this study, we investigate the Post-Soviet cultural food knowledge of foundation models by constructing BORSch, a multimodal dataset encompassing 1147 and 823 dishes in the Russian and Ukrainian languages, centered around the Post-Soviet region. We demonstrate that leading models struggle to correctly identify the origins of dishes from Post-Soviet nations in both text-only and multimodal Question Answering (QA), instead over-predicting countries linked to the language the question is asked in. Through analysis of pretraining data, we show that these results can be explained by misleading dish-origin co-occurrences, along with linguistic phenomena such as Russian-Ukrainian code mixing. Finally, to move beyond QA-based assessments, we test models' abilities to produce accurate visual descriptions of dishes. The weak correlation between this task and QA suggests that QA alone may be insufficient as an evaluation of cultural understanding. To foster further research, we will make BORSch publicly available at https://github.com/alavrouk/BORSch.

文化理解多模态偏见分析饮食知识

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。