arXiv:2606.07541cs.HCcs.AI2026-06中稿 · SocialLLM @ ICWSM …

用大模型模拟人类看视频的主观感受,发现效果有限但有改进空间。

Multimodal Large Language Models as Synthetic Participants in Video-Based Studies: An Evaluation

论文配图:Multimodal Large Language Models as Synthetic Participants in Video-Based Studies: An Evaluation
图 1 · 摘自论文原文
  • 基于感知刺激价值框架,用大模型生成虚拟参与者评分
  • 顶尖模型与真人评分相关性低,存在系统性偏差和群体差异失真
  • 提示策略影响复杂,部分改善但可能恶化其他方面,适合研究人机交互者

多模态大语言模型(MLLM)在视频理解等客观任务上表现优异,但其能否模拟依赖社会背景的主观人类反应仍不明确。为此,我们评估了MLLM作为视频研究中合成参与者的有效性,聚焦于短视频的感官沉浸感感知评估。基于感知消息刺激价值(PMSV)框架,使用17项量表比较673名真实参与者与受个人特征条件控制的MLLM模拟评分,涵盖情绪唤醒、戏剧冲击力和新奇性。结果表明,即使领先模型(Gemini 3 Flash 和 Qwen 3 Omni)与真人评分一致性有限,表现出显著的均值下偏和集中趋势偏差,扭曲了子群体差异,并对用户画像敏感性不一致。不同提示策略对各项指标影响各异,仅小幅提升部分维度而恶化另一些。这些发现揭示了将MLLM用于视频研究合成参与者时的挑战与潜力。数据与代码:https://github.com/MINDLab25/mllm-human-simulation-eval

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have shown strong performance on objective tasks such as video understanding and reasoning. However, it remains unclear whether they can approximate subjective human responses, which depend not only on content comprehension but also on individuals' social contexts. To address this gap, we evaluate MLLMs as synthetic participants in an emerging task: assessing perceived sensory engagement with short videos. Grounded in the Perceived Message Sensation Value (PMSV) framework, we compare ratings from recruited human participants and profile-conditioned MLLM simulations (n=673) using a 17-item scale measuring emotional arousal, dramatic impact, and novelty. We find that even leading MLLMs (Gemini 3 Flash and Qwen 3 Omni) show limited agreement with human participants. The models exhibit distinct downward mean-shift and central-tendency biases in their rating distributions. They both introduce and flatten subgroup differences, while showing inconsistent sensitivity to participant profiles. Prompting strategies affect these metrics differently, modestly improving some aspects while worsening others. These results highlight both the challenges and opportunities of developing MLLMs as synthetic participants in video-based research. Data and code: https://github.com/MINDLab25/mllm-human-simulation-eval

多模态模型主观评估合成数据人机模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。