arXiv:2510.02343cs.CLcs.AI2025-10被引 7

构建隐私保护的社交媒体用户数据集,用于训练和评估大模型的社会角色模拟能力。

$\texttt{BluePrint}$: A Social Media User Dataset for LLM Persona Evaluation and Training

  • 通过行为聚类与匿名化技术构建可复用的社交用户数据集
  • 包含12种互动类型,支持上下文依赖的社交行为建模
  • 适用于政治话语、虚假信息等社会问题的研究

大语言模型(LLMs)在大规模模拟社交媒体动态方面展现出巨大潜力,但缺乏用于微调和评估模型作为真实社交代理的标准数据资源。本文提出SIMPACT(SIMulation-oriented Persona and Action Capture Toolkit),一个尊重隐私的框架,用于构建基于行为的社交媒体数据集。将下一步行为预测设为训练与评估任务,并引入群体与聚类层面的评估指标,衡量行为保真度与风格真实性。作为具体实现,我们发布BluePrint——基于公开Bluesky数据构建的大规模政治话语数据集。该数据集对匿名用户进行行为聚类,形成代表聚合行为的用户角色,通过伪名化和移除个人标识信息保障隐私。数据包含12种社交互动类型(点赞、回复、转发等),每条记录关联前序发帖行为。该数据集支持构建考虑语言与交互行为双重上下文的代理模型。通过标准化数据与评估协议,SIMPACT为推进严谨、伦理合规的社交媒体模拟提供基础。BluePrint既可用作政治话语建模的评估基准,也可作为构建特定领域数据集的模板,以研究虚假信息、极化等挑战。

原文摘要 · Abstract (English)

Large language models (LLMs) offer promising capabilities for simulating social media dynamics at scale, enabling studies that would be ethically or logistically challenging with human subjects. However, the field lacks standardized data resources for fine-tuning and evaluating LLMs as realistic social media agents. We address this gap by introducing SIMPACT, the SIMulation-oriented Persona and Action Capture Toolkit, a privacy respecting framework for constructing behaviorally-grounded social media datasets suitable for training agent models. We formulate next-action prediction as a task for training and evaluating LLM-based agents and introduce metrics at both the cluster and population levels to assess behavioral fidelity and stylistic realism. As a concrete implementation, we release BluePrint, a large-scale dataset built from public Bluesky data focused on political discourse. BluePrint clusters anonymized users into personas of aggregated behaviours, capturing authentic engagement patterns while safeguarding privacy through pseudonymization and removal of personally identifiable information. The dataset includes a sizable action set of 12 social media interaction types (likes, replies, reposts, etc.), each instance tied to the posting activity preceding it. This supports the development of agents that use context-dependence, not only in the language, but also in the interaction behaviours of social media to model social media users. By standardizing data and evaluation protocols, SIMPACT provides a foundation for advancing rigorous, ethically responsible social media simulations. BluePrint serves as both an evaluation benchmark for political discourse modeling and a template for building domain specific datasets to study challenges such as misinformation and polarization.

大模型社交模拟数据集隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。