用大模型生成跨平台高保真社交数据,解决真实数据难获取问题。
Towards High-Fidelity Synthetic Multi-platform Social Media Datasets via Large Language Models
- 基于多平台话题提示,用大模型生成跨平台社交帖子。
- 不同模型生成效果差异明显,需后处理提升数据真实性。
- 提出新评估指标,适合研究虚假信息与跨平台传播。
社交平台数据对虚假信息、影响力操作、仇恨言论检测及意见领袖营销等研究至关重要,但因成本和平台限制,获取困难,尤其是跨平台数据更难。本文探索大语言模型生成跨平台、语义与词汇相关的合成社交数据的潜力,旨在达到真实数据质量。我们提出多平台话题引导策略,使用多个语言模型基于两个真实数据集(每组含三个平台的帖子)生成合成数据,并评估其词汇与语义特性,与真实数据对比。实证结果表明,大模型生成跨平台数据具有前景,不同模型表现不一,且需后处理以提升保真度。除对三种前沿大模型的评估外,本研究还提出针对多平台社交数据的新保真度评估指标。
原文摘要 · Abstract (English)
Social media datasets are essential for research on a variety of topics, such as disinformation, influence operations, hate speech detection, or influencer marketing practices. However, access to social media datasets is often constrained due to costs and platform restrictions. Acquiring datasets that span multiple platforms, which is crucial for understanding the digital ecosystem, is particularly challenging. This paper explores the potential of large language models to create lexically and semantically relevant social media datasets across multiple platforms, aiming to match the quality of real data. We propose multi-platform topic-based prompting and employ various language models to generate synthetic data from two real datasets, each consisting of posts from three different social media platforms. We assess the lexical and semantic properties of the synthetic data and compare them with those of the real data. Our empirical findings show that using large language models to generate synthetic multi-platform social media data is promising, different language models perform differently in terms of fidelity, and a post-processing approach might be needed for generating high-fidelity synthetic datasets for research. In addition to the empirical evaluation of three state of the art large language models, our contributions include new fidelity metrics specific to multi-platform social media datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。