arXiv:2510.11563cs.CL2025-10被引 7

首个评估大模型跨文化对话能力的框架与基准,揭示当前模型在文化适应上的显著不足。

Culturally-Aware Conversations: A Framework & Benchmark for LLMs

  • 基于社会文化理论构建对话风格与情境、关系、文化背景的关联框架
  • 通过多元文化评审者标注数据集,发现主流大模型在跨文化对话中表现不佳
  • 提出对话框架、风格敏感性等新评估标准,适合关注AI文化适配的研究者

现有评估大模型文化适应能力的基准与真实跨文化对话挑战脱节。本文提出首个面向真实多文化对话场景的框架与基准。该框架基于社会文化理论,形式化语言风格这一文化沟通关键要素如何受情境、关系和文化背景影响。我们据此构建了由多元文化评审者标注的基准数据集,并提出适用于NLP跨文化评估的新标准:对话框架、风格敏感性和主观正确性。对当前顶级大模型的评估显示,这些模型在对话场景中的文化适应能力存在明显短板。

原文摘要 · Abstract (English)

Existing benchmarks that measure cultural adaptation in LLMs are misaligned with the actual challenges these models face when interacting with users from diverse cultural backgrounds. In this work, we introduce the first framework and benchmark designed to evaluate LLMs in realistic, multicultural conversational settings. Grounded in sociocultural theory, our framework formalizes how linguistic style - a key element of cultural communication - is shaped by situational, relational, and cultural context. We construct a benchmark dataset based on this framework, annotated by culturally diverse raters, and propose a new set of desiderata for cross-cultural evaluation in NLP: conversational framing, stylistic sensitivity, and subjective correctness. We evaluate today's top LLMs on our benchmark and show that these models struggle with cultural adaptation in a conversational setting.

大模型跨文化对话系统评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。