arXiv:2603.20231cs.CYcs.AI2026-03

LLMs在职场沟通中更懂分寸,能提升邮件质量。

Moral Mazes in the Era of LLMs

  • 用游戏化任务测试人类与LLM写职场邮件的风格差异。
  • 人类邮件通过LLM重写后通过率超50%,表现优于原生文本。
  • 更强的AI判官偏好含蓄表达,暗示沟通规范正在演变。

职场中应对复杂社交情境至关重要,如提批评意见不伤士气、拒绝对方请求不疏远同事。尽管大语言模型(LLMs)正渗透工作场景,但其在处理此类规范方面的表现尚不明确。为此,我们构建了HR Simulator——一款让用户扮演人力资源专员,撰写应对挑战性职场情境的邮件,并由GPT-4o作为裁判,依据情境特定评分标准进行评估。分析超过600封人类与LLM撰写的邮件发现,LLM邮件更正式且更具同理心。此外,人类邮件表现逊于LLM(如23.5%对48-54%的情境通过率),但经由LLM重写的邮件可超越两者,显示混合优势。在评估层面,不同判官模型存在偏好差异:10个裁判模型的分析表明,较弱模型偏爱直接表达,而更强模型倾向更微妙的措辞;随着模型能力提升,判官间一致性增强,暗示向共享沟通规范趋同,可能与人类习惯不同。整体结果表明,若广泛采用,LLMs可能显著重塑职场沟通方式。

原文摘要 · Abstract (English)

Navigating complex social situations is an integral part of corporate life, ranging from giving critical feedback without hurting morale to rejecting requests without alienating teammates. Although large language models (LLMs) are permeating the workplace, it is unclear how well they can navigate these norms. To investigate this question, we created HR Simulator, a game where users roleplay as an HR officer and write emails to tackle challenging workplace scenarios, evaluated with GPT-4o as a judge based on scenario-specific rubrics. We analyze over 600 human and LLM emails and find systematic differences in style: LLM emails are more formal and empathetic. Furthermore, humans underperform LLMs (e.g., 23.5% vs. 48-54% scenario pass rate), but human emails rewritten by LLMs can outperform both, which indicates a hybrid advantage. On the evaluation side, judges can exhibit differences in their email preferences: an analysis of 10 judge models reveals evidence for emergent tact, where weaker models prefer direct, blunt communication but stronger models prefer more subtle messages. Judges also agree with each other more as they scale, which hints at a convergence toward shared communicative norms that may differ from humans'. Overall, our results suggest LLMs could substantially reshape communication in the workplace if they are widely adopted in professional correspondence.

职场沟通LLM评估生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。