构建真实人机互动数据集,评测大模型社会推理能力。
Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning
- 构建400段真实人机交互视频数据集,含超1万条标注。
- 大模型在识别社交错误与合理回应上表现远低于人类。
- 适合研究人机协作、社交智能的AI开发者使用。
本文聚焦基础模型在真实人机交互中的社会推理能力,提出社会人机具身对话(SHREC)数据集,包含约400段真实世界人机交互视频及超过1万条标注,涵盖机器人社交失误、能力表现、底层原因和修正方案。与以往关注人与人互动的数据集不同,该数据集突出真实社交机器人面临的情感理解、意图追踪和对话机制等挑战。当前基础模型难以识别这些细微且情境相关的失败。为此,我们定义了八个基准任务,覆盖社交错误与能力检测、社会属性识别、交互流程理解及理由与正确行为生成。对最先进基础模型的实验与人工评估显示显著性能差距,凸显该任务难度,并为发展更具社会智能的AI指明方向。
原文摘要 · Abstract (English)
Our work focuses on the social reasoning capabilities of foundation models for real-world human-robot interactions. We introduce the Social Human Robot Embodied Conversation (SHREC) Dataset, a benchmark of $\sim$400 real-world human-robot interaction videos and over 10K annotations, capturing robot social errors, competencies, underlying rationales, and corrections. Unlike prior datasets focused on human-human interactions, the SHREC Dataset uniquely highlights the social challenges faced by real-world social robots such as emotion understanding, intention tracking, and conversational mechanics. Moreover, current foundation models struggle to recognize these deficits, which manifest as subtle, socially situated failures. To evaluate AI models' capacity for social reasoning, we define eight benchmark tasks targeting critical areas such as (1) detection of social errors and competencies, (2) identification of underlying social attributes, (3) comprehension of interaction flow, and (4) providing rationale and alternative correct actions. Experiments with state-of-the-art foundation models, alongside human evaluations, reveal substantial performance gaps -- underscoring the difficulty and providing directions in developing socially intelligent AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。