arXiv:2606.03284cs.CL2026-06

构建首个面向东南亚文化的自然语言推理数据集,揭示大模型在跨文化理解上的短板。

SEA-NLI: Natural Language Inference as a Lens into Southeast Asian Cultural Understanding

论文配图:SEA-NLI: Natural Language Inference as a Lens into Southeast Asian Cultural Understanding
图 1 · 摘自论文原文
  • 构建覆盖8国的原生多语言推理数据集,由母语者验证真实性
  • 所有模型在东南亚文化相关任务上表现差,知识类题目准确率不足40%
  • 引入文化适配与提示增强可提升性能,思维链提示效果有限

前沿大模型在西方语境下表现良好,但在东南亚(SEA)等代表性不足的文化中仍缺乏充分测试。现有自然语言推理(NLI)基准大多以西方为中心、基于翻译或仅限单语,难以衡量文化相关的推理能力。我们提出SEA-NLI,一个原生、文化根植的NLI基准,涵盖8个东南亚国家的英语及本土语言,经母语者验证。对17种编码器与解码器模型的评估显示,所有模型表现普遍偏低,尤其在语言与科技类知识密集型任务上。分析表明,失败主要源于缺乏东南亚文化知识:使用适应东南亚文化的模型和文化感知提示能显著提升表现,而思维链(CoT)提示带来的增益有限。

原文摘要 · Abstract (English)

Frontier LLMs perform well in Western contexts, but remain poorly tested on underrepresented cultures such as those in Southeast Asia (SEA). Existing NLI benchmarks are largely Western-centric, translation-derived, or monolingual, limiting their ability to measure culturally grounded reasoning. We introduce SEA-NLI, a native, culturally grounded NLI benchmark covering eight SEA countries in English and native regional languages, verified by native speakers. Across 17 encoder and decoder models, we observe a low performance from all models, especially for knowledge-intensive categories such as Languages and Science and Technology. Our analysis shows that failure cases mainly stem from missing SEA cultural knowledge: SEA-adapted models and culture-aware prompting improve performance, while CoT prompting offers limited gains.

自然语言推理跨文化理解多语言文化知识

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。