arXiv:2604.28048cs.CLcs.SI2026-04被引 2

用虚拟人格测试大模型对城市情绪的判断,发现效果有限且易偏激。

Stable Behavior, Limited Variation: Persona Validity in LLM Agents for Urban Sentiment Perception

论文配图:Stable Behavior, Limited Variation: Persona Validity in LLM Agents for Urban Sentiment Perception
图 1 · 摘自论文原文
  • 给大模型分配不同虚拟人格,测试其对城市图像的情绪判断差异。
  • 同人格下判断稳定,但跨人格差异小,经济与性格影响微弱。
  • 去掉人格提示后模型反而更接近人类判断,说明人格标签作用不大。

大型语言模型(LLMs)正被用于模拟城市感知中的人类判断,但虚拟人格提示是否能带来有意义且可复现的行为多样性仍不明确。本文通过涵盖性别、经济状况、政治倾向和性格特征的因子化人格设置,构建多代理系统评估来自PerceptSent数据集的城市场景图像,并分析同一人格内的稳定性与跨人格间的差异性。结果显示,共享同一人格的代理间表现出强一致性,行为稳定可复现;然而跨人格差异有限:经济状况与性格虽有统计显著性但实际影响微小,性别无明显效应,政治倾向仅产生可忽略的影响。此外,代理存在极端化偏差,导致人类标注中常见的中间情绪类别被压缩。因此,在粗粒度情感极性任务上表现良好,但随着情感分辨精度提升而性能下降,表明简单的标签式人格提示无法捕捉精细感知判断。进一步对比无虚拟人格条件下的同一模型,发现其在多数任务变体中与人类标注的一致性甚至优于带人格提示的版本,暗示在该场景下简单标签式人格提示对标注价值贡献有限。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used as proxies for human perception in urban analysis, yet it remains unclear whether persona prompting produces meaningful and reproducible behavioral diversity. We investigate whether distinct personas influence urban sentiment judgments generated by multimodal LLMs. Using a factorial set of personas spanning gender, economic status, political orientation, and personality, we instantiate multiple agents per persona to evaluate urban scene images from the PerceptSent dataset and assess both within-persona consistency and cross-persona variation. Results show strong convergence among agents sharing a persona, indicating stable and reproducible behavior. However, cross-persona differentiation is limited: economic status and personality induce statistically detectable but practically modest variation, while gender shows no measurable effect and political orientation only negligible impact. Agents also exhibit an extremity bias, collapsing intermediate sentiment categories common in human annotations. As a result, performance remains strong on coarse-grained polarity tasks but degrades as sentiment resolution increases, suggesting that simple label-based persona prompting does not capture fine-grained perceptual judgments. To isolate the contribution of persona conditioning, we additionally evaluate the same model without personas. Surprisingly, the no-persona model sometimes matches or exceeds persona-conditioned agreement with human labels across all task variants, suggesting that simple label-based persona prompting may add limited annotation value in this setting.

大模型城市感知人格建模情绪识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。