arXiv:2608.28405cs.CLcs.CY2026-08

构建多轮文化对话评估框架,提升大模型在东亚东南亚的跨文化助人能力。

CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia

论文配图:CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia
图 1 · 摘自论文原文
  • 设计多轮仿真系统,模拟10个地区、58个亚群体的文化场景对话。
  • 包含14,610个评测对话与27万条黄金对话,支持跨域能力验证。
  • 微调2.7万条高质量数据可提升文化理解与安全判断能力,适合本地化应用研究者。

当前大语言模型的文化评估多依赖单轮事实问答,难以反映用户在真实文化情境中寻求持续帮助的常见需求。本文提出CultureConverse,一个可扩展的多语言仿真与评估工具,覆盖东亚及东南亚10个地区、58个子群体身份和7个应用领域。每个模拟对话生成带评分的交互记录,助手需在部分信息下推断文化约束并提供协助。由此构建的CultureConverse-DS数据集包含14,610个基准(评估)对话和274,295条由人工标注引导的黄金模式对话。对18个模型的基准评估显示,GPT-5 mini表现最优。人类标注实验表明该评估框架可有效替代人工评判。基于27,860条高质量样本微调后,模型在领域内协助质量提升,并在文化选择题与安全分类任务中实现跨域迁移效果。本文发布仿真工具、两个数据集划分及评价提示,支持交互式文化能力评估。

原文摘要 · Abstract (English)

Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common use case: users seeking practical help over multiple turns in culturally grounded scenarios. We introduce CultureConverse, a scalable, multilingual simulation and evaluation harness for culturally grounded assistant dialogue that covers 10 East and Southeast Asian regions, 58 subgroup identities, and 7 domains. Each simulated and evaluated episode produces a scored interaction where the assistant assists the user and infers cultural constraints from partial information. The resulting CultureConverse-DS dataset contains 14,610 benchmark (evaluation) episodes and 274,295 oracle-guided (gold-mode) dialogues. In our benchmark evaluation of 18 models, GPT-5 mini achieves the highest assistance quality. Human annotation experiments suggest that our evaluation framework is a sufficient proxy for human judgment. Performance gains from fine-tuning on 27,860 high-quality CultureConverse-DS samples improve in-domain assistance and transfer out-of-domain to cultural MCQ and safety classification benchmarks. We release the harness, both splits, and judge prompts to support interactive evaluation of cultural competency.

文化智能多轮对话评估框架大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。