让大模型具备显式共情能力,才能真正理解人类情感与立场。
LLMs Should Incorporate Explicit Mechanisms for Human Empathy
- 将共情定义为可观察的行为属性,识别出四大共情失败机制。
- 实证发现模型高分表现下仍存在情感弱化、语义失真等系统性问题。
- 适用于需理解人类情绪与关系的高风险场景,如医疗、心理辅导。
本文主张大型语言模型(LLMs)应引入显式的共情机制。随着LLMs在高风险人本场景中广泛应用,其成功不仅取决于正确性或流畅性,更在于能否忠实保留人类视角。然而,当前的LLMs在这一要求上系统性失败:即便对齐良好且符合政策,仍常弱化情感、误判情境重要性、僵化关系立场,导致意义扭曲。我们把共情定义为可观测的行为属性——即建模并回应人类视角,同时保持意图、情感与上下文。在此框架下,我们识别出当代LLMs中的四种共情失败机制:情感弱化、共情粒度错配、冲突回避和语言疏离,这些是现有训练与对齐实践的结构性后果。我们进一步从认知、文化、关系三个维度组织这些失败,解释其在任务中的表现。实证分析表明,强基准表现可能掩盖系统性共情失真,因此亟需将共情感知的目标、评测与训练信号作为大模型开发的核心组成部分。
原文摘要 · Abstract (English)
This paper argues that Large Language Models (LLMs) should incorporate explicit mechanisms for human empathy. As LLMs become increasingly deployed in high-stakes human-centered settings, their success depends not only on correctness or fluency but on faithful preservation of human perspectives. Yet, current LLMs systematically fail at this requirement: even when well-aligned and policy-compliant, they often attenuate affect, misrepresent contextual salience, and rigidify relational stance in ways that distort meaning. We formalize empathy as an observable behavioral property: the capacity to model and respond to human perspectives while preserving intention, affect, and context. Under this framing, we identify four recurring mechanisms of empathic failure in contemporary LLMs--sentiment attenuation, empathic granularity mismatch, conflict avoidance, and linguistic distancing--arising as structural consequences of prevailing training and alignment practices. We further organize these failures along three dimensions: cognitive, cultural, and relational empathy, to explain their manifestation across tasks. Empirical analyses show that strong benchmark performance can mask systematic empathic distortions, motivating empathy-aware objectives, benchmarks, and training signals as first-class components of LLM development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。