arXiv:2602.03107cs.CL2026-02被引 1

评测大模型对中文礼貌、不礼貌和伪礼貌的理解能力,发现性能差异显著。

The Mask of Civility: Benchmarking Chinese Mock Politeness Comprehension in Large Language Models

  • 用礼貌管理理论构建三类中文语用数据集。
  • 六款模型在不同提示下表现差异大,知识增强策略提升明显。
  • 适合研究语言智能与人文社科交叉的学者参考。

从语用学视角出发,本研究系统评估了代表性大语言模型在识别中文礼貌、不礼貌及伪礼貌现象上的表现差异。针对现有语用理解空白,研究基于关系管理理论与伪礼貌模型,构建了一个融合真实与模拟话语的三分类数据集。选取GPT-5.1、DeepSeek等六款代表性模型,在零样本、少样本、知识增强及混合策略四种提示条件下进行评测。该研究是‘宏大语言学’范式下的有益尝试,为技术变革时代语用理论的应用提供新路径,回应了技术与人文学科如何共存的当代议题,是一次语言技术与人文反思的跨学科探索。

原文摘要 · Abstract (English)

From a pragmatic perspective, this study systematically evaluates the differences in performance among representative large language models (LLMs) in recognizing politeness, impoliteness, and mock politeness phenomena in Chinese. Addressing the existing gaps in pragmatic comprehension, the research adopts the frameworks of Rapport Management Theory and the Model of Mock Politeness to construct a three-category dataset combining authentic and simulated Chinese discourse. Six representative models, including GPT-5.1 and DeepSeek, were selected as test subjects and evaluated under four prompting conditions: zero-shot, few-shot, knowledge-enhanced, and hybrid strategies. This study serves as a meaningful attempt within the paradigm of ``Great Linguistics,'' offering a novel approach to applying pragmatic theory in the age of technological transformation. It also responds to the contemporary question of how technology and the humanities may coexist, representing an interdisciplinary endeavor that bridges linguistic technology and humanistic reflection.

语用理解大模型评测中文对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。