arXiv:2607.21016cs.CL2026-07

首个印尼本土语言对话式文化常识评测集,助力大模型理解真实语境中的文化差异。

CultureTalk-ID: A Multi-Task Dialogue Benchmark for Cultural Commonsense in Indonesian Local Languages

论文配图:CultureTalk-ID: A Multi-Task Dialogue Benchmark for Cultural Commonsense in Indonesian Local Languages
图 1 · 摘自论文原文
  • 构建11种印尼方言的4496段对话,覆盖13个文化议题
  • 提出三项任务:文化常识推理、文化保真翻译与语言引导生成
  • 由母语者多阶段筛选,确保文化真实性,适合多语言文化研究

文化通过对话得以体现,但现有印尼文化常识评测仅基于孤立短句,丢失了文化细节实际浮现的对话语境。我们提出CultureTalk-ID,首个基于对话的文化常识评测集,涵盖11种印尼本地语言和13个文化敏感主题,共4,496段由母语者参与的多阶段人工标注对话,确保真实性。该评测引入三项互补任务:基于对话的多选文化常识推理、文化保真机器翻译、语言引导生成,共同检验大模型是否具备理解、迁移与生成文化语境化语言的能力。

原文摘要 · Abstract (English)

Culture is lived through conversation, yet existing Indonesian cultural commonsense benchmarks evaluate LLMs on short and isolated prompts, stripping away the dialogic context in which cultural nuances actually surface. We introduce CultureTalk-ID, the first dialogue-based benchmark for cultural commonsense in Indonesian and its local languages, comprising 4,496 culturally grounded dialogues across 11 languages and 13 culturally salient topics, curated through a multi-stage human pipeline with native speakers to ensure authenticity. CultureTalk-ID introduces three complementary tasks, namely dialogue-based multiple-choice cultural commonsense reasoning, culturally faithful machine translation, and language steering, which jointly probe whether LLMs can understand, transfer, and generate culturally grounded language.

文化常识多语言对话评测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。