arXiv:2602.02975cs.CLcs.AI2026-02AAAI被引 1

测试大模型在情境中理解社会规范的能力,发现其表现仍有明显短板。

Where Norms and References Collide: Evaluating LLMs on Normative Reasoning

  • 构建情境化规范测试集SNIC,评估模型对隐含规范的理解能力。
  • 强模型在隐性或冲突规范下识别准确率不足,难以稳定推理。
  • 适合研究具身智能、社会规范推理的学者参考。

具身智能体(如机器人)需在具体环境中互动,成功沟通常依赖于对社会规范的推理:即在特定情境中约束行为适当性的共享预期。关键能力之一是基于规范的指称消解(NBRR),其要求理解指称表达时推断出物理与社会情境中的隐含规范性预期。然而,当前大型语言模型(LLMs)是否具备此类推理能力尚不明确。本文提出SNIC(Situated Norms in Context),一个经人工验证的情境化诊断测试平台,用于探测先进LLMs在NBRR任务中提取和运用相关规范原则的能力。SNIC聚焦于日常任务(如清洁、整理、服务)中产生的物理基础规范。在一系列控制实验中,我们发现即使是最先进的模型也难以一致地识别并应用社会规范,尤其当规范隐含、表述不明确或相互冲突时。这些结果揭示了当前大模型的一个认知盲区,凸显了将语言系统部署于社会性具身场景中的核心挑战。

原文摘要 · Abstract (English)

Embodied agents, such as robots, will need to interact in situated environments where successful communication often depends on reasoning over social norms: shared expectations that constrain what actions are appropriate in context. A key capability in such settings is norm-based reference resolution (NBRR), where interpreting referential expressions requires inferring implicit normative expectations grounded in physical and social context. Yet it remains unclear whether Large Language Models (LLMs) can support this kind of reasoning. In this work, we introduce SNIC (Situated Norms in Context), a human-validated diagnostic testbed designed to probe how well state-of-the-art LLMs can extract and utilize normative principles relevant to NBRR. SNIC emphasizes physically grounded norms that arise in everyday tasks such as cleaning, tidying, and serving. Across a range of controlled evaluations, we find that even the strongest LLMs struggle to consistently identify and apply social norms, particularly when norms are implicit, underspecified, or in conflict. These findings reveal a blind spot in current LLMs and highlight a key challenge for deploying language-based systems in socially situated, embodied settings.

规范推理具身智能语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。