评估多模态大模型在图文场景下的社会规范推理能力,发现文本优于图像。
Social Norm Reasoning in Multimodal Language Models: An Evaluation
- 用30个图文故事测试5个多模态模型的规范推理能力
- GPT-4o在图文任务中表现最佳,文本准确率显著高于图像
- 复杂规范推理仍是所有模型的共同难点,适合机器人社交系统研究者参考
在多智能体系统(MAS)中,智能体需具备理解与推理社会规范的能力以实现人际交互(如人机协作)。规范性多智能体系统(NorMAS)研究规范的生成、违规检测与惩罚机制。现有研究多采用符号化方法(如形式逻辑)表示规范,仅适用于简化环境。相比之下,多模态大语言模型(MLLMs)有望在包含文本和图像的复杂社会情境中实现规范识别与推理。然而,此前研究局限于纯文本场景。本文通过30个文本故事与30个图像故事,评估5个MLLM在规范相关问题上的表现,并与人类答案对比。结果表明,MLLMs在文本场景中的规范推理能力优于图像场景。GPT-4o在两种模态下均表现最佳,展现了与多智能体系统集成的潜力;免费模型Qwen-2.5VL次之。所有模型对复杂规范的推理仍存在困难。
原文摘要 · Abstract (English)
In Multi-Agent Systems (MAS), agents are designed with social capabilities, allowing them to understand and reason about social concepts such as norms when interacting with others (e.g., inter-robot interactions). In Normative MAS (NorMAS), researchers study how norms develop, and how violations are detected and sanctioned. However, existing research in NorMAS use symbolic approaches (e.g., formal logic) for norm representation and reasoning whose application is limited to simplified environments. In contrast, Multimodal Large Language Models (MLLMs) present promising possibilities to develop software used by robots to identify and reason about norms in a wide variety of complex social situations embodied in text and images. However, prior work on norm reasoning have been limited to text-based scenarios. This paper investigates the norm reasoning competence of five MLLMs by evaluating their ability to answer norm-related questions based on thirty text-based and thirty image-based stories, and comparing their responses against humans. Our results show that MLLMs demonstrate superior performance in norm reasoning in text than in images. GPT-4o performs the best in both modalities offering the most promise for integration with MAS, followed by the free model Qwen-2.5VL. Additionally, all models find reasoning about complex norms challenging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。