arXiv:2512.14720cs.SIcs.AI2025-12AAAI被引 6

首个评估大模型社交代理的现实基准,涵盖海量真实社交数据。

SoMe: A Realistic Benchmark for LLM-based Social Media Agents

  • 构建包含超900万帖子的多任务评测平台,模拟真实社交环境。
  • 主流大模型在复杂社交任务中表现不佳,闭源与开源模型均存短板。
  • 适合研究社交智能体、大模型评估及数字人文方向的研究者使用。

由大语言模型驱动的智能代理在社交媒体平台展现出惊人能力并日益流行。然而,当前缺乏对这类代理理解媒体内容、分析用户行为和做出复杂决策能力的全面评估。为此,我们提出SoMe——首个面向大模型社交代理的综合性基准,支持多种代理工具访问与分析社交数据。该基准包含8类社交任务、9,164,284条帖子、6,591个用户档案、25,686份外部报告,以及17,869条精细标注的任务查询。相比现有数据集与基准,SoMe首次提供了一个多样化且真实的测试环境。通过大量定量与定性分析,我们首次揭示主流智能体大模型在真实社交场景中的性能表现,并识别出若干局限。评估显示,当前闭源与开源大模型均无法令人满意地完成社交代理任务。SoMe为未来社交智能体的发展提供了具有挑战性且有意义的测试平台。代码与数据已公开于 https://github.com/LivXue/SoMe。

原文摘要 · Abstract (English)

Intelligent agents powered by large language models (LLMs) have recently demonstrated impressive capabilities and gained increasing popularity on social media platforms. While LLM agents are reshaping the ecology of social media, there exists a current gap in conducting a comprehensive evaluation of their ability to comprehend media content, understand user behaviors, and make intricate decisions. To address this challenge, we introduce SoMe, a pioneering benchmark designed to evaluate social media agents equipped with various agent tools for accessing and analyzing social media data. SoMe comprises a diverse collection of 8 social media agent tasks, 9,164,284 posts, 6,591 user profiles, and 25,686 reports from various social media platforms and external websites, with 17,869 meticulously annotated task queries. Compared with the existing datasets and benchmarks for social media tasks, SoMe is the first to provide a versatile and realistic platform for LLM-based social media agents to handle diverse social media tasks. By extensive quantitative and qualitative analysis, we provide the first overview insight into the performance of mainstream agentic LLMs in realistic social media environments and identify several limitations. Our evaluation reveals that both the current closed-source and open-source LLMs cannot handle social media agent tasks satisfactorily. SoMe provides a challenging yet meaningful testbed for future social media agents. Our code and data are available at https://github.com/LivXue/SoMe

大模型评估社交智能体基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。