用人格提示影响城市视觉理解中的语言生成,发现解释比描述更受人格影响。
Persona Prompting in Multimodal Urban Perception: Descriptive Convergence and Interpretive Variation

- 按描述、标签、解释三层分析模型输出,探究人格提示作用
- 经济地位差异导致解释差异最大,描述基本一致
- 适合研究人机交互中主观认知差异的学者参考
本研究探讨人格提示如何影响两个多模态大模型在城市感知场景下的语言生成,该场景用于分析对共享视觉证据的主观解读。我们将输出分为三类:描述性基础(图像描述)、中间语义层(感知标签)和解释性框架(理由说明)。基于Qwen3-VL-8B与Gemma-4-E4B-it模型,每模型生成约60,000条人格条件化标注。结果显示,不同人格下的描述高度趋同,仅在属性相关细节上存在微小差异;而理由说明则显著分化,经济地位影响最大,政治倾向与个性亦明显。图像级配对比较证实,这三项属性在理由上的差异大于描述。感知标签方面,相同属性层级的人格产生更相似的标签集,经济地位区分最明显。探索性主题分析揭示人格特有的评价侧重。所有输出类型中,人格对之间的相似性模式在模型间高度相关,但理由的一致性最低。总体而言,人格提示对解释性框架的影响强于描述性基础。
原文摘要 · Abstract (English)
This study examines how persona prompting shapes language generated by two multimodal large language models in urban perception, a setting for examining subjective interpretations of shared visual evidence. We organize outputs into three functional levels: descriptive grounding (captions), intermediate semantic layer (perception tags), and interpretive framing (justifications). Using approximately 60,000 persona-conditioned annotations per model from Qwen3-VL-8B and Gemma-4-E4B-it, we find that captions converge strongly across persona profiles and show only small attribute-associated differences. Justifications vary substantially more: economic status produces the largest difference in both models, with political orientation and personality also prominent. Paired image-level comparisons confirm larger justification than caption differences for these three attributes. For perception tags, personas sharing the same attribute level produce more similar tag sets than personas with different attribute levels, with the largest separation observed for economic status. Exploratory topic analysis further reveals persona-specific evaluative emphasis. Across models, profile-pair similarity patterns are strongly correlated for all three output types, although agreement is lowest for justifications. Overall, persona prompting affects interpretive framing more strongly than descriptive grounding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。