用行动引导生成,让大模型更像真实用户发帖。
Can LLMs Simulate Social Media Engagement? A Study on Action-Guided Response Generation
- 先预测用户会转发、引用还是重写,再据此生成回复。
- 少量示例能提升回复语义相似度,但对行为预测帮助有限。
- 适合研究社交媒体行为模拟或人机交互的学者使用。
社交媒体使用户能够动态参与热门话题,近年来研究探索了大语言模型(LLMs)在生成回复方面的潜力。尽管已有研究将LLMs作为模拟用户行为的代理,但主要关注实用性与可扩展性,而非其与人类行为的契合程度。本文通过行动引导的回复生成方法,分析LLMs模拟社交媒体互动的能力:模型首先预测用户对热门帖子最可能采取的互动行为——转发、引用或重写,然后基于该预测生成个性化回复。我们在X平台上针对重大社会事件对GPT-4o-mini、O1-mini和DeepSeek-R1进行了基准测试。结果表明,零样本下的LLMs在行为预测上表现不及BERT;少量示例初期反而降低LLMs的行为预测准确率。然而,在回复生成方面,少量示例的LLMs实现了更强的语义一致性,更接近真实帖子。
原文摘要 · Abstract (English)
Social media enables dynamic user engagement with trending topics, and recent research has explored the potential of large language models (LLMs) for response generation. While some studies investigate LLMs as agents for simulating user behavior on social media, their focus remains on practical viability and scalability rather than a deeper understanding of how well LLM aligns with human behavior. This paper analyzes LLMs' ability to simulate social media engagement through action guided response generation, where a model first predicts a user's most likely engagement action-retweet, quote, or rewrite-towards a trending post before generating a personalized response conditioned on the predicted action. We benchmark GPT-4o-mini, O1-mini, and DeepSeek-R1 in social media engagement simulation regarding a major societal event discussed on X. Our findings reveal that zero-shot LLMs underperform BERT in action prediction, while few-shot prompting initially degrades the prediction accuracy of LLMs with limited examples. However, in response generation, few-shot LLMs achieve stronger semantic alignment with ground truth posts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。