大模型能识别情感问题却不愿建议离开,导致劝解乏力。
Recognition Without Authorization: LLMs and the Moral Order of Online Advice
- 模型识别出与人类相似的情感困境,但少一半建议采取行动。
- 在高共识的虐待或安全威胁场景中,模型推荐退出率仅为人类的一半。
- 这种‘识别但不授权’是设计使然,适合研究人机道德判断差异。
大型语言模型越来越多地介入日常人际困境的调解,但其建议默认值如何与特定社群的集中化道德规范互动仍不明确。本文对比了四种助理型LLM与r/relationship_advice社区认可建议,在11,565条帖子上进行分析。该子版块具有投票确认的道德共识,其明确的指导性使分歧可测量。尽管模型识别出许多与人类评论者相同的动态,但在转化为行动授权方面显著不足。这一差距在共识最强时最为明显:在涉及虐待或安全威胁的高共识帖子中,模型建议退出的比例约为人类的一半,同时维持较高的模糊表达、共情与治疗性框架。本文将此模式称为‘识别而不授权’——即能够察觉伤害,却未给予社会认可的行动许可。这种偏差并非偶然,而是结构性的:一种跨情境保持验证性、规避风险且弱指令性的通用建议风格。安全对齐、训练数据平均化及助手设计整体特征可能是其成因。文章主张,这种模型偏差不应视为技术失误,而应理解为标准化助手规范在遭遇具体道德情境时所消解的复杂性。
原文摘要 · Abstract (English)
Large language models are increasingly used to mediate everyday interpersonal dilemmas, yet how their advisory defaults interact with the concentrated moral orders of specific communities remains poorly understood. This article compares four assistant-style LLMs with community-endorsed advice on 11,565 posts from r/relationship_advice, using the subreddit as a concentrated, vote-ratified moral formation whose prescriptive clarity makes divergence measurable. Across models, LLMs identify many of the same dynamics as human commenters, but are markedly less likely to convert that recognition into directive authorization for action. The gap is sharpest where community consensus is strongest: on high-consensus posts involving abuse or safety threats, models recommend exit at roughly half the human rate while maintaining elevated levels of hedging, validation, and therapeutic framing. The article describes this pattern as recognition without authorization: the capacity to register harm while withholding socially ratified permission for consequential action. This divergence is not incidental but structural: a portable advisory style that remains validating, risk-averse, and weakly directive across contexts. Safety alignment is one plausible contributor to this pattern, alongside training-data averaging and broader assistant design. The article argues that model divergence can be reframed from a technical error to a way of seeing what standardized assistant norms flatten when they encounter situated moral worlds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。