arXiv:2608.07498cs.HCcs.AI2026-08

大模型能精准预测社交用户反应,但依赖完整个人资料。

Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation

  • 用12种大模型基于人物档案预测点赞/点踩,测试不同资料完整度影响。
  • 全资料下最高准确率达96.68%,仅用人口统计信息时降至51.00%。
  • 模型表现受选择和推理能力影响,适合用于推荐系统压力测试。

社交平台中的自主AI代理对民主话语和平台治理构成现实风险,同时也可作为推荐系统部署前的测试工具。核心问题是:基于人物设定的LLM能否以足够精度模拟个体社交媒体反应?该研究在296个基于调查的人物档案和26个真实标注帖子上,对12种LLM配置进行二元点赞/点踩预测基准测试,涵盖三种资料完整度条件,并以留帖外机器学习分类器为基线。全资料条件下准确率范围为75.54%至96.68%,30个百分点差异主要由模型选择导致,经配对McNemar检验与代理级自助区间验证。GPT-5.5 Pro在全资料下准确率达96.68%,减资料时降至62.32%,仅含人口统计信息时跌至51.00%,与多数类基线无显著差异,证实仅靠人口统计无法提供有效预测信号。监督分类器在留帖外测试中降至15.4%,而LLM保持真正的零样本泛化能力。自适应推理提升部分模型表现。有直接资料锚点的帖子,模型间一致性(平均κ=0.44)是无锚点帖子(κ=0.23)的近两倍;最同质配置使34%的模拟群体反应趋于一致。结果验证了基于LLM的仿真可用于推荐系统压力测试,同时揭示了大规模合成代理群对公共舆论构成可信威胁的行为准确性。

原文摘要 · Abstract (English)

Autonomous AI agents in social media present concrete risks to democratic discourse and platform governance, while also offering tools for pre-deployment recommender system testing. A central open question is whether persona-prompted LLMs can simulate individual-level social media reactions with sufficient accuracy to support either application, and how accuracy depends on profile completeness, model selection, and the generalization challenge posed by novel post content. This study benchmarks twelve LLM configurations on binary like/dislike prediction across 296 survey-based agent profiles and 26 ground-truth-mapped posts under three profile conditions, with leave-post-out machine learning classifiers as baselines. Across full-profile conditions, accuracy ranges from 75.54% to 96.68%, with a 30-point spread attributable primarily to model selection and confirmed by paired McNemar tests with agent-level bootstrap intervals. GPT-5.5 Pro accuracy degrades monotonically from 96.68% under a full profile to 62.32% under a reduced profile and to 51.00% with demographics alone, the last indistinguishable from the majority-class baseline, which confirms that demographic inference provides negligible predictive signal. Supervised classifiers collapse to 15.4% under leave-post-out, while LLMs sustain genuine zero-shot generalization unavailable to trained methods. Adaptive reasoning improves accuracy substantially for some models. Inter-model agreement is nearly double for posts with direct profile anchors (mean \k{appa} = 0.44) than for posts without them (\k{appa} = 0.23), and the least heterogeneous configuration homogenizes 34% of simulated population reactions. Results validate LLM-based simulation for recommender system stress-testing while documenting the behavioral accuracy that makes large-scale synthetic agent swarms a credible threat to public opinion.

大模型社交模拟推荐系统行为预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。