构建真实对话基准,评估大模型社交智能表现
SI-Bench: Benchmarking Social Intelligence of Large Language Models in Human-to-Human Conversations
- 基于真实社交应用对话数据构建评测集
- 顶尖模型在复杂情境推理上超人类,但回复质量仍不足
- 思维链提示反而降低社交对话表现,适合研究人机交互者
随着大语言模型(LLMs)日益具备类人能力,它们被越来越多地用作与人类交互的自主代理。然而,评估其在真实复杂社交互动中的表现仍是重大挑战。以往研究多依赖模拟的代理间对话构建数据集,难以捕捉真实人类对话的语言风格与关系动态。为此,我们提出SI-Bench,一个基于社会科学研究理论的新基准,包含从社交应用中收集的2,221条真实多轮对话,并对其中312条对话进行8个主流模型的人工标注。实验表明,当前最优模型在复杂社交情境下的过程推理能力已超过人类专家,但在回复质量上仍落后于人类。此外,引入思维链(CoT)推理反而会损害大模型在社交对话任务中的表现。所有数据集已公开于https://github.com/SI-Bench/SI-Bench.git。
原文摘要 · Abstract (English)
As large language models (LLMs) develop anthropomorphic abilities, they are increasingly being deployed as autonomous agents to interact with humans. However, evaluating their performance in realistic and complex social interactions remains a significant challenge. Most previous research built datasets through simulated agent-to-agent interactions, which fails to capture the authentic linguistic styles and relational dynamics found in real human conversations. To address this gap, we introduce SI-Bench, a novel benchmark designed to evaluate aspects of social intelligence in LLMs. Grounded in broad social science theories, SI-Bench contains 2,221 authentic multi-turn dialogues collected from a social networking application. We further selected a subset of 312 dialogues for manual annotation across 8 major models. The experiments show that SOTA models have surpassed the human expert in process reasoning under complex social situations, yet they still fall behind humans in reply quality. Moreover, introducing Chain-of-Thought (CoT) reasoning may degrade the performance of LLMs in social dialogue tasks. All datasets are openly available at https://github.com/SI-Bench/SI-Bench.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。