测试大模型在群聊中识别隐性社交规范的能力,发现多数模型表现有限。
LoSoNA: A Benchmark for Local Social Norm Adaptation in Group Conversations

- 通过群聊对话片段让模型推断隐藏的社交规范
- 显式提示下部分模型准确率最高达84.2%
- 适合评估大模型社会交互能力的研究者使用
在线群聊具有未明说的本地对话规范,基于大模型的智能体识别并适应这些规范的能力尚未被充分探索。本文提出LoSoNA,一个用于多参与方聊天中局部社交规范适应性的基准测试。每个场景提供一段精心设计的群聊记录,其中非目标参与者表现出隐藏的本地规范,随后以引导性发言迫使目标模型作答,以揭示其是否推断出该规范。我们在四种提示条件下评估了八个前沿及开源模型。朴素提示对大多数模型效果不佳;明确提示虽有提升,但效果不均:Gemini 3.1 Pro 达到84.2%准确率,Claude Fable 5 达到81.6%,其余部分模型增益微弱甚至退步。该基准响应近期对大模型社交能力评估的呼吁,检验模型能否从先前对话中推断本地规范,并在单轮回复中应用。
原文摘要 · Abstract (English)
Online group chats are social spaces with local conversational norms that are rarely stated explicitly. The ability and willingness of LLM-based agents to recognize and adapt to these norms remains mostly unexplored. We introduce LoSoNA, a benchmark for local social norm adaptation in multi-party chat. Each scenario gives a subject model a curated group-chat transcript in which non-subject participants demonstrate a hidden local norm, followed by a final elicitor turn that forces a response revealing whether the subject has inferred that norm. We evaluate eight frontier and open-weight models under four prompting conditions that vary how explicitly the model is told to treat the prior conversation as evidence for how it should answer. Naive prompting remains limited for most models; explicit norm-aware prompting helps unevenly, with Gemini 3.1 Pro reaching $84.2\%$ and Claude Fable 5 reaching $81.6\%$, while several other models show small gains or regressions. LoSoNA contributes to recent calls for evaluating LLM social capabilities by testing whether models can infer local conversational norms from precedent and use them in a one-turn group-chat response.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。