arXiv:2607.02049cs.CLcs.AI2026-07

评测大模型在英乌双语情感支持中的表现,发现语言切换影响回应质量。

SPLIT: Cross-Lingual Empathy and Cultural Grounding in English and Ukrainian LLM Responses

论文配图:SPLIT: Cross-Lingual Empathy and Cultural Grounding in English and Ukrainian LLM Responses
图 1 · 摘自论文原文
  • 构建500提示的跨语言情感评估基准SPLIT,覆盖五类心理困境场景。
  • 谷歌Gemini与LLaMA在乌克兰语中表现下降,DeepSeek-V3相对稳定。
  • 人类与AI对文化贴合度评价差异大,凸显人工评估的重要性。

大型语言模型在情感支持和危机情境中日益应用,但其跨语言能力仍待深入研究。现有基准多关注多语言性能,却很少考察低资源语言中的情感共情与文化适配性。本文提出SPLIT,一个包含500个提示的基准,用于评估模型在压力、恐慌、孤独、内部流离、紧张五类情境下生成情绪化回应的一致性。我们评估了三种技术路径不同的大模型,从共情准确性、语言自然度、上下文与文化适配性三个维度进行比较。结果表明,Gemini-2.5-Flash和LLaMA-3.3-70B-Instruct在转向乌克兰语时性能显著下降,而DeepSeek-V3保持相对稳定。此外,人类与AI评估者在共情与自然度上仅有弱相关,但在文化适配性上分歧明显。研究进一步指出,生成乌克兰语文本不等于提供有效的乌克兰情感支持。这些发现有助于未来设计更贴近文化的评估体系,并推动以用户为中心的评价机制发展。

原文摘要 · Abstract (English)

Large Language Models are increasingly deployed in emotional-support contexts and crisis-related situations. Nevertheless, their cross-lingual abilities in these circumstances remain underexplored. Existing benchmarks emphasize multilingual performance but rarely examine crisis-related empathy and cultural grounding in low-to-mid-resource languages. We introduce SPLIT, a 500-prompt benchmark designed to evaluate LLM consistency in generating emotionally grounded responses across five categories: Stress, Panic, Loneliness, Internal Displacement, and Tension. We evaluate three technically diverse LLMs across three dimensions: Empathetic Accuracy, Linguistic Naturalness, and Contextual & Cultural Grounding. The framework aims to assess and compare the quality of LLM responses in both English and Ukrainian languages, as well as to explore the reliability of the LLM-as-a-jury paradigm. Our findings reveal that Gemini-2.5-Flash and LLaMA-3.3-70B-Instruct degrade when transitioning to Ukrainian, while DeepSeek-V3 remains comparatively stable within our benchmark. We additionally find that human and AI evaluators agree weakly on empathy and naturalness but diverge on cultural grounding. We further argue that producing Ukrainian text is not equivalent to producing Ukrainian emotional support. Our findings may assist in the future development of more culturally tailored benchmark designs, as well as encourage a stronger emphasis on human-centered evaluation.

情感计算跨语言文化适配大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。