arXiv:2505.04628cs.CLcs.AI2025-05被引 4

构建首个评估大模型多用户社交能力的基准,测其在复杂互动中的表现。

How Social is It? A Benchmark for LLMs' Capabilities in Multi-user Multi-turn Social Agent Tasks

  • 基于社会学框架设计四阶段任务流程,模拟真实社交场景。
  • 实验表明大模型在角色切换与持续对话中表现有限,依赖思维链提升但耗能高。
  • 提出新指标衡量思维链效率,平衡准确率与计算成本,适合评测社交型AI。

将大语言模型(LLMs)从单一用户辅助角色拓展至复杂社会场景中的多用户、多轮社交代理,亟需系统性评估其社会能力。现有基准尚不完善。为此,本文提出基于社会学原理的代理任务分级框架,并构建全新基准HSII(How Social Is It),用于评估大模型在综合社交任务中的表现。该基准包含四个阶段:格式解析、目标选择、目标切换对话与稳定对话,依托从新闻数据集逐步构建的HSII-Dataset。通过聚类分析对数据进行结构化处理,并研究思维链(Chain of Thought, COT)对社交性能的影响。由于COT计算开销大,本文引入新统计指标“COT复杂度”,量化不同模型在特定社交任务中使用COT的效率,实现正确性与效率的权衡。实验结果表明,该基准能有效评估大模型的社会能力。

原文摘要 · Abstract (English)

Expanding the application of large language models (LLMs) to societal life, instead of primary function only as auxiliary assistants to communicate with only one person at a time, necessitates LLMs' capabilities to independently play roles in multi-user, multi-turn social agent tasks within complex social settings. However, currently the capability has not been systematically measured with available benchmarks. To address this gap, we first introduce an agent task leveling framework grounded in sociological principles. Concurrently, we propose a novel benchmark, How Social Is It (we call it HSII below), designed to assess LLM's social capabilities in comprehensive social agents tasks and benchmark representative models. HSII comprises four stages: format parsing, target selection, target switching conversation, and stable conversation, which collectively evaluate the communication and task completion capabilities of LLMs within realistic social interaction scenarios dataset, HSII-Dataset. The dataset is derived step by step from news dataset. We perform an ablation study by doing clustering to the dataset. Additionally, we investigate the impact of chain of thought (COT) method on enhancing LLMs' social performance. Since COT cost more computation, we further introduce a new statistical metric, COT-complexity, to quantify the efficiency of certain LLMs with COTs for specific social tasks and strike a better trade-off between measurement of correctness and efficiency. Various results of our experiments demonstrate that our benchmark is well-suited for evaluating social skills in LLMs.

大模型社交能力基准测试思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。