arXiv:2604.18955cs.CLcs.AI2026-04

评测大模型在社交媒体分析中的多任务能力,涵盖身份验证、内容生成和用户属性推断。

Assessing Capabilities of Large Language Models in Social Media Analytics: A Multi-task Quest

论文配图:Assessing Capabilities of Large Language Models in Social Media Analytics: A Multi-task Quest
图 1 · 摘自论文原文
  • 构建统一评估框架,覆盖作者身份验证、内容生成与用户属性推断三任务。
  • 在2024年1月后新收集数据上验证泛化能力,减少已见数据偏差。
  • 通过真实用户感知实验,揭示大模型生成内容的可信度与局限性。

本研究首次系统评估了GPT-4、GPT-4o、GPT-3.5-Turbo、Gemini 1.5 Pro、DeepSeek-V3、Llama 3.2和BERT等现代大模型在推特(X)数据集上的三大核心社交媒体分析任务:(I)作者身份验证,(II)帖子生成,(III)用户属性推断。针对身份验证,提出基于多样化用户与帖子采样策略的系统性框架,并在2024年1月后的新增推文上测试泛化能力,以缓解“已见数据”偏差。对于帖子生成,采用综合评估指标衡量大模型生成符合用户风格的真实内容的能力。衔接任务一与二,开展用户研究,评估真实用户对基于自身写作风格生成内容的感知效果。针对属性推断,使用IAB Tech Lab 2023和2018 U.S. SOC两个标准化分类体系标注职业与兴趣,并与现有基线对比。整体评估提供新见解并建立可复现基准。代码与数据将随论文发布。

原文摘要 · Abstract (English)

In this study, we present the first comprehensive evaluation of modern LLMs - including GPT-4, GPT-4o, GPT-3.5-Turbo, Gemini 1.5 Pro, DeepSeek-V3, Llama 3.2, and BERT - across three core social media analytics tasks on a Twitter (X) dataset: (I) Social Media Authorship Verification, (II) Social Media Post Generation, and (III) User Attribute Inference. For the authorship verification, we introduce a systematic sampling framework over diverse user and post selection strategies and evaluate generalization on newly collected tweets from January 2024 onward to mitigate "seen-data" bias. For post generation, we assess the ability of LLMs to produce authentic, user-like content using comprehensive evaluation metrics. Bridging Tasks I and II, we conduct a user study to measure real users' perceptions of LLM-generated posts conditioned on their own writing. For attribute inference, we annotate occupations and interests using two standardized taxonomies (IAB Tech Lab 2023 and 2018 U.S. SOC) and benchmark LLMs against existing baselines. Overall, our unified evaluation provides new insights and establishes reproducible benchmarks for LLM-driven social media analytics. The code and data are provided in the supplementary material and will also be made publicly available upon publication.

大模型评估社交媒体分析作者验证内容生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。