arXiv:2605.17079cs.CLcs.AI2026-05被引 1

测试大模型能否像真实消费者一样预测舆论反应,发现主流模型表现远未达标。

Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench

论文配图:Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench
图 1 · 摘自论文原文
  • 用1553个真实社交话题和2.3万条可审计的反应点构建评测基准
  • 最强模型仅覆盖47.8%的真实反应标准,多数模型严重滞后
  • 揭示技术强弱与社会认知之间的巨大鸿沟,适合评估消费模拟能力

大模型被越来越多地用作‘数字消费者’来模拟公众意见、预测试营销决策并预测受众反应。然而,现有评估很少检验模型是否能重建真实消费者在公共话语中表现出的具体反应模式。本文提出ConsumerSimBench,基于1,553个真实中文社交媒体话题和23,122个经过规则审核的原子级反应标准,涵盖四类反应类型。不同于依赖整体偏好判断的开放式生成评分,该基准将每个任务分解为可审计的二元判断,使三名评审员一致性从65.8%提升至92.1%,且点对点判断与人类多数意见一致率达98.4%。在13个前沿生成模型中,最强模型Gemini-3.1-Pro仅覆盖47.8%的真实反应标准,GPT-5.2和Claude-4.6虽在技术基准上表现优异,但在本任务中大幅落后。失败揭示了技术性能与社会情境理解之间的显著差距。直接使用结构化推理提示反而降低覆盖率,而采用generate–reflect多智能体流程将MiMo-V2.5-Pro的覆盖率从32.9%提升至37.6%。ConsumerSimBench将消费者模拟重构为对真实公共话语反应的预测问题,表明当前前沿大模型仍远无法可靠预测中国高语境消费话语中消费者真正关注的内容。

原文摘要 · Abstract (English)

LLMs are increasingly used as ``digital consumers'' to simulate public opinion, pre-test marketing decisions, and anticipate audience response. However, existing evaluations rarely ask whether a model can reconstruct the concrete reaction patterns that real consumers surface in public discourse. We introduce ConsumerSimBench, a benchmark built from 1,553 real Chinese social-media topics and 23,122 atomic, rule-audited criteria spanning four reaction families. Rather than scoring open-ended generations with a holistic preference judge, ConsumerSimBench decomposes each task into auditable yes-no decisions over concrete reaction points, raising three-judge agreement from 65.8% to 92.1% with 98.4% agreement between pointwise judge decisions and human-majority labels. Across 13 frontier generators, the strongest model, Gemini-3.1-Pro, covers only 47.8% of real reaction criteria, while GPT-5.2 and Claude-4.6 trail far behind despite their strength on technical benchmarks. The failures reveal a sharp gap between technical-benchmark performance and socially grounded consumer intuition. A direct structured reasoning prompt decreases coverage, while a generate--reflect multi-agent pipeline improves MiMo-V2.5-Pro from 32.9% to 37.6% on a subset. ConsumerSimBench reframes consumer simulation as a forecasting problem over real public-discourse reactions, showing that frontier LLMs remain far from reliably predicting what consumers will actually care about in high-context Chinese consumer discourse.

大模型评估消费者模拟社会反应中文语境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。