arXiv:2602.22752cs.CLcs.AI2026-02中稿 · the 15th Workshop …被引 5

用LLM模拟社交媒体用户,验证其行为预测的可靠性。

Towards Simulating Social Media Users with LLMs: Evaluating the Operational Validity of Conditioned Comment Prediction

  • 通过对比生成评论与真实数字痕迹,评估LLM模拟用户行为能力。
  • 微调使文本结构更像真人,但削弱了内容语义一致性。
  • 直接分析用户历史比虚构个人简介更能提升模拟精度。

大型语言模型(LLMs)从探索性工具转向社会科学研究中的‘硅基主体’,但其操作有效性缺乏充分验证。本研究提出条件评论预测(CCP)任务:模型根据给定刺激预测用户会如何评论,并将生成结果与真实数字痕迹对比。该框架系统评估了当前LLM在模拟社交媒体用户行为方面的能力。我们在英语、德语和卢森堡语场景下测试了开源8B参数模型(Llama3.1、Qwen3、Ministral)。通过对比显式与隐式提示策略及监督微调(SFT)的影响,发现低资源环境下存在形式与内容的解耦现象:虽然微调使输出长度和语法结构更接近真实,但削弱了语义根基。此外,我们证明在微调后,显式条件化(生成人物背景)变得冗余,因为模型能直接从行为历史中进行潜在推断。研究挑战了现有的‘天真提示’范式,提出应优先使用真实行为痕迹而非描述性人格设定,以实现高保真度模拟。

原文摘要 · Abstract (English)

The transition of Large Language Models (LLMs) from exploratory tools to active "silicon subjects" in social science lacks extensive validation of operational validity. This study introduces Conditioned Comment Prediction (CCP), a task in which a model predicts how a user would comment on a given stimulus by comparing generated outputs with authentic digital traces. This framework enables a rigorous evaluation of current LLM capabilities with respect to the simulation of social media user behavior. We evaluated open-weight 8B models (Llama3.1, Qwen3, Ministral) in English, German, and Luxembourgish language scenarios. By systematically comparing prompting strategies (explicit vs. implicit) and the impact of Supervised Fine-Tuning (SFT), we identify a critical form vs. content decoupling in low-resource settings: while SFT aligns the surface structure of the text output (length and syntax), it degrades semantic grounding. Furthermore, we demonstrate that explicit conditioning (generated biographies) becomes redundant under fine-tuning, as models successfully perform latent inference directly from behavioral histories. Our findings challenge current "naive prompting" paradigms and offer operational guidelines prioritizing authentic behavioral traces over descriptive personas for high-fidelity simulation.

LLM模拟社交媒体行为预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。