让大模型学会判断行为是否合乎社会规范,提升与人类共识的对齐度。
SocialGaze: Improving the Integration of Human Social Norms in Large Language Models
- 通过多视角推理框架让模型先描述社会情境再判断
- 相比原模型,对齐人类判断的准确率提升最高达11个F1点
- 发现模型对性别和年龄存在偏见,男性更易被不公平归责
尽管近年来大量研究致力于提升大语言模型(LLMs)的推理能力,但其与社会价值和规范的对齐仍存在明显差距。本文提出‘社会接受度判断’任务,要求模型评估特定社交情境下行为的可接受性,例如邻居夜间要求社区成员将宠物留在室内是否合理。实验发现,当前大模型在社会接受度理解上常偏离人类共识。为此,我们提出SocialGaze——一种多步提示框架,使模型从多个视角解析社会情境后再形成判断。结果表明,该方法可使GPT-3.5模型在社会判断上与人类一致性的F1值提升最高达11点。此外,我们还识别出模型在归责时存在偏差:男性角色更易被不公平地归责,而对年长叙述者的判断更接近人类共识。
原文摘要 · Abstract (English)
While much research has explored enhancing the reasoning capabilities of large language models (LLMs) in the last few years, there is a gap in understanding the alignment of these models with social values and norms. We introduce the task of judging social acceptance. Social acceptance requires models to judge and rationalize the acceptability of people's actions in social situations. For example, is it socially acceptable for a neighbor to ask others in the community to keep their pets indoors at night? We find that LLMs' understanding of social acceptance is often misaligned with human consensus. To alleviate this, we introduce SocialGaze, a multi-step prompting framework, in which a language model verbalizes a social situation from multiple perspectives before forming a judgment. Our experiments demonstrate that the SocialGaze approach improves the alignment with human judgments by up to 11 F1 points with the GPT-3.5 model. We also identify biases and correlations in LLMs in assigning blame that is related to features such as the gender (males are significantly more likely to be judged unfairly) and age (LLMs are more aligned with humans for older narrators).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。