arXiv:2508.07069cs.CLcs.AI2025-08被引 3

构建东南亚多语言对话数据集,融入文化背景提升对话真实感

SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages

  • 聚焦东南亚8种低资源语言,包含人物属性与本土话题
  • 涵盖6国多轮对话,每轮对话嵌入文化情境与生活细节
  • 适合研究跨文化对话系统、个性化聊天机器人开发者

尽管已有大量对话数据集支持对话系统研究,但多数闲聊数据集忽略了自然对话中的文化差异。为弥补这一空白,我们推出SEADialogues,一个以东南亚为中心的文化语境对话数据集。该地区拥有超过7亿人口和丰富的文化多样性。数据集包含来自六个国家的八种语言对话,许多语言虽使用者众多却属低资源语言。每段对话均包含人物特征及两个反映当地日常生活的文化相关话题,以增强对话的文化相关性与个性化。此外,我们发布一个多轮对话数据集,旨在推动具备文化意识与以人为本特性的大语言模型研究,包括对话代理系统的开发。

原文摘要 · Abstract (English)

Although numerous datasets have been developed to support dialogue systems, most existing chit-chat datasets overlook the cultural nuances inherent in natural human conversations. To address this gap, we introduce SEADialogues, a culturally grounded dialogue dataset centered on Southeast Asia, a region with over 700 million people and immense cultural diversity. Our dataset features dialogues in eight languages from six Southeast Asian countries, many of which are low-resource despite having sizable speaker populations. To enhance cultural relevance and personalization, each dialogue includes persona attributes and two culturally grounded topics that reflect everyday life in the respective communities. Furthermore, we release a multi-turn dialogue dataset to advance research on culturally aware and human-centric large language models, including conversational dialogue agents.

多语言对话文化建模低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。