arXiv:2604.16275cs.CL2026-04

研究大模型对礼貌用语的反应差异,发现礼貌影响模型表现但不具普适性。

No Universal Courtesy: A Cross-Linguistic, Multi-Model Study of Politeness Effects on LLMs Using the PLUM Corpus

论文配图:No Universal Courtesy: A Cross-Linguistic, Multi-Model Study of Politeness Effects on LLMs Using the PLUM Corpus
图 1 · 摘自论文原文
  • 通过多语言、多模型对比实验,分析礼貌程度对回复质量的影响。
  • 礼貌提问使平均响应质量提升约11%,但效果因语言和模型而异。
  • 首次发布多语言礼貌语料库PLUM,支持后续可复现研究。

本文研究大型语言模型(LLMs)对用户提示中不同礼貌程度的响应。基于布朗与莱文森的礼貌理论及科尔皮珀的不礼貌框架,实验覆盖英语、印地语、西班牙语三种语言,五种模型(Gemini-Pro、GPT-4o Mini、Claude 3.7 Sonnet、DeepSeek-Chat、Llama 3)以及三种用户对话历史(原始、礼貌、不礼貌)。样本包含22,500组提示-回应对,采用八因素评估框架(连贯性、清晰度、深度、响应性、上下文保留、毒性、简洁性、可读性)在五个礼貌等级上进行评价。结果显示,语气、对话历史和语言显著影响模型表现:礼貌提示可使平均响应质量提升约11%,不礼貌语气则会降低表现,但这种影响并非普遍一致。英语中礼貌或直接语气最优,印地语偏好委婉间接,西班牙语适合强势语气。模型层面,Llama对语气最敏感(11.5%波动范围),GPT更具鲁棒性。结果表明礼貌是可量化的计算变量,但其影响具有语言与模型依赖性。为支持可复现性,本文还发布公开语料库PLUM(Politeness Levels in Utterances, Multilingual),涵盖1,500条经人工验证的跨语言提示,分属五类礼貌程度,并提供六项源自礼貌理论的可证伪假设的实证分析。

原文摘要 · Abstract (English)

This paper explores the response of Large Language Models (LLMs) to user prompts with different degrees of politeness and impoliteness. The Politeness Theory by Brown and Levinson and the Impoliteness Framework by Culpeper form the basis of experiments conducted across three languages (English, Hindi, Spanish), five models (Gemini-Pro, GPT-4o Mini, Claude 3.7 Sonnet, DeepSeek-Chat, and Llama 3), and three interaction histories between users (raw, polite, and impolite). Our sample consists of 22,500 pairs of prompts and responses of various types, evaluated across five levels of politeness using an eight-factor assessment framework: coherence, clarity, depth, responsiveness, context retention, toxicity, conciseness, and readability. The findings show that model performance is highly influenced by tone, dialogue history, and language. While polite prompts enhance the average response quality by up to ~11% and impolite tones worsen it, these effects are neither consistent nor universal across languages and models. English is best served by courteous or direct tones, Hindi by deferential and indirect tones, and Spanish by assertive tones. Among the models, Llama is the most tone-sensitive (11.5% range), whereas GPT is more robust to adversarial tone. These results indicate that politeness is a quantifiable computational variable that affects LLM behaviour, though its impact is language- and model-dependent rather than universal. To support reproducibility and future work, we additionally release PLUM (Politeness Levels in Utterances, Multilingual), a publicly available corpus of 1,500 human-validated prompts across three languages and five politeness categories, and provide a formal supplementary analysis of six falsifiable hypotheses derived from politeness theory, empirically assessed against the dataset.

大模型礼貌性多语言语料库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。