arXiv:2506.09391cs.CL2025-06EMNLP被引 10

对比人类与大模型在自由表达中的礼貌策略,发现大模型更偏消极礼貌但更受欢迎。

Comparing human and LLM politeness strategies in free production

  • 比较人类与大模型在受限和开放任务中的礼貌表达方式。
  • 700亿参数以上模型复现了文献中的礼貌偏好,且在开放场景中更受好评。
  • 模型过度使用消极策略,易导致语义误读,提示需关注对话的语用对齐。

礼貌言语是大型语言模型(LLMs)面临的核心对齐挑战。人类会灵活运用语言策略,在传递信息与维护社交关系间取得平衡——包括建立亲和力的正面策略(如赞美、表达兴趣),以及减少冒犯的负面策略(如缓和语气、间接表达)。我们通过对比人类与大模型在受限和开放式生成任务中的回应,探究模型是否具备类似的情境敏感性。研究发现,参数量≥700亿的模型能有效复现计算语用学文献中的关键礼貌偏好;在开放式语境下,人类评估者反而更偏好模型生成的内容。然而进一步的语言分析表明,即便在积极情境中,模型仍过度依赖负向礼貌策略,可能引发误解。尽管现代大模型已表现出对礼貌策略的出色掌握,这些细微差异凸显了人工智能系统在语用对齐方面的深层问题。

原文摘要 · Abstract (English)

Polite speech poses a fundamental alignment challenge for large language models (LLMs). Humans deploy a rich repertoire of linguistic strategies to balance informational and social goals -- from positive approaches that build rapport (compliments, expressions of interest) to negative strategies that minimize imposition (hedging, indirectness). We investigate whether LLMs employ a similarly context-sensitive repertoire by comparing human and LLM responses in both constrained and open-ended production tasks. We find that larger models ($\ge$70B parameters) successfully replicate key preferences from the computational pragmatics literature, and human evaluators surprisingly prefer LLM-generated responses in open-ended contexts. However, further linguistic analyses reveal that models disproportionately rely on negative politeness strategies even in positive contexts, potentially leading to misinterpretations. While modern LLMs demonstrate an impressive handle on politeness strategies, these subtle differences raise important questions about pragmatic alignment in AI systems.

礼貌策略大模型语用对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。