构建新基准评估智能体在真实场景中协调物理与社交约束的能力
S$^3$IT: A Benchmark for Spatially Situated Social Intelligence Test
- 设计座位排序任务,让智能体在3D环境中协调多人偏好与关系
- 现有大模型在该任务上表现远低于人类,暴露空间智能短板
- 适合研究具身智能、人机交互与多智能体协作的学者
将具身智能体融入人类环境需要同时理解社交规范与物理限制。现有评估要么仅关注文本中的抽象社交推理,要么忽略社交因素的物理任务,均无法检验智能体在真实具身情境中平衡物理与社交约束的能力。为此,我们提出空间定位社交智能测试(S$^{3}$IT),聚焦一个新颖且具有挑战性的座位排序任务:在3D环境中为由大语言模型驱动的多个具有不同身份、偏好和复杂人际关系的非玩家角色(NPC)安排座位。我们的可程序扩展框架生成多样且可控难度的场景,要求智能体通过主动对话获取偏好、自主探索感知环境,并在复杂约束网络中进行多目标优化。我们在S$^{3}$IT上评估了当前最先进大模型,发现其表现显著落后于人类基线,表明大模型在空间智能方面存在明显缺陷,但对具有明确文本线索的冲突解决仍可达到接近人类水平的表现。
原文摘要 · Abstract (English)
The integration of embodied agents into human environments demands embodied social intelligence: reasoning over both social norms and physical constraints. However, existing evaluations fail to address this integration, as they are limited to either disembodied social reasoning (e.g., in text) or socially-agnostic physical tasks. Both approaches fail to assess an agent's ability to integrate and trade off both physical and social constraints within a realistic, embodied context. To address this challenge, we introduce Spatially Situated Social Intelligence Test (S$^{3}$IT), a benchmark specifically designed to evaluate embodied social intelligence. It is centered on a novel and challenging seat-ordering task, requiring an agent to arrange seating in a 3D environment for a group of large language model-driven (LLM-driven) NPCs with diverse identities, preferences, and intricate interpersonal relationships. Our procedurally extensible framework generates a vast and diverse scenario space with controllable difficulty, compelling the agent to acquire preferences through active dialogue, perceive the environment via autonomous exploration, and perform multi-objective optimization within a complex constraint network. We evaluate state-of-the-art LLMs on S$^{3}$IT and found that they still struggle with this problem, showing an obvious gap compared with the human baseline. Results imply that LLMs have deficiencies in spatial intelligence, yet simultaneously demonstrate their ability to achieve near human-level competence in resolving conflicts that possess explicit textual cues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。