arXiv:2601.02830cs.CL2026-01

对比中美大模型在中文文化任务上的表现差异

The performances of the Chinese and U.S. Large Language Models on the Topic of Chinese Culture

  • 采用直接提问法评估多款中美的大模型对中华文化理解能力
  • 中国模型整体表现优于美国模型,尤其在历史与诗歌领域
  • 性能差异或源于训练数据分布与文化内容重视程度不同

文化背景影响个体的视角与问题解决方式。自2018年GPT-1问世以来,大语言模型(LLMs)发展迅速。目前全球十大领先LLM开发者主要集中在中国和美国。为检验由中国和美国开发者发布的模型在中文语境下是否存在文化差异,本文针对中国传统文化相关问题评估了GPT-5.1、DeepSeek-V3.2、Qwen3-Max和Gemini2.5Pro等模型的表现。评估涵盖历史、文学、诗歌等传统中国文化领域。对比分析显示,中国开发的模型在这些任务上普遍表现更优。在美国开发的模型中,Gemini 2.5Pro和GPT-5.1表现相对较高。观察到的性能差异可能源于训练数据分布、本地化策略及模型开发过程中对中国文化内容的重视程度不同。

原文摘要 · Abstract (English)

Cultural backgrounds shape individuals' perspectives and approaches to problem-solving. Since the emergence of GPT-1 in 2018, large language models (LLMs) have undergone rapid development. To date, the world's ten leading LLM developers are primarily based in China and the United States. To examine whether LLMs released by Chinese and U.S. developers exhibit cultural differences in Chinese-language settings, we evaluate their performance on questions about Chinese culture. This study adopts a direct-questioning paradigm to evaluate models such as GPT-5.1, DeepSeek-V3.2, Qwen3-Max, and Gemini2.5Pro. We assess their understanding of traditional Chinese culture, including history, literature, poetry, and related domains. Comparative analyses between LLMs developed in China and the U.S. indicate that Chinese models generally outperform their U.S. counterparts on these tasks. Among U.S.-developed models, Gemini 2.5Pro and GPT-5.1 achieve relatively higher accuracy. The observed performance differences may potentially arise from variations in training data distribution, localization strategies, and the degree of emphasis on Chinese cultural content during model development.

大模型评估文化差异中文理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。