对比图文表述对大模型人格表现的影响,发现图像更一致但易被忽略。
A Thousand Words or An Image: Studying the Influence of Persona Modality in Multimodal LLMs
- 构建四模态人格数据集,含图文、文字、图文组合及视觉化文字
- 5个多模态模型测试显示,文字描述更显语言习惯,视觉文字更贴合人格
- 揭示模型常忽略图像中的人格细节,适合研究多模态人格建模的学者
大型语言模型(LLM)在模拟多样人格方面取得显著进展,提升了对话代理和虚拟助手的效能。尽管人类人格可通过文本或图像表达,但其模态对模型人格呈现的影响仍不明确。本文构建了一个包含40种不同年龄、性别、职业与地域人格的平行模态数据集,涵盖四种表达方式:仅图像、仅文本、图像+小段文本、以及通过排版设计传达属性的视觉文字。我们设计了包含60个问题的评估框架与指标,系统评估5个多模态LLM在不同属性与场景下的人格表现。实验表明,详细文本表示能更好体现语言习惯,而视觉文字更保持一致性;但模型普遍忽视图像中的特定人格信息,暴露出当前技术局限,为后续研究指明方向。数据与代码已开源。
原文摘要 · Abstract (English)
Large language models (LLMs) have recently demonstrated remarkable advancements in embodying diverse personas, enhancing their effectiveness as conversational agents and virtual assistants. Consequently, LLMs have made significant strides in processing and integrating multimodal information. However, even though human personas can be expressed in both text and image, the extent to which the modality of a persona impacts the embodiment by the LLM remains largely unexplored. In this paper, we investigate how do different modalities influence the expressiveness of personas in multimodal LLMs. To this end, we create a novel modality-parallel dataset of 40 diverse personas varying in age, gender, occupation, and location. This consists of four modalities to equivalently represent a persona: image-only, text-only, a combination of image and small text, and typographical images, where text is visually stylized to convey persona-related attributes. We then create a systematic evaluation framework with 60 questions and corresponding metrics to assess how well LLMs embody each persona across its attributes and scenarios. Comprehensive experiments on $5$ multimodal LLMs show that personas represented by detailed text show more linguistic habits, while typographical images often show more consistency with the persona. Our results reveal that LLMs often overlook persona-specific details conveyed through images, highlighting underlying limitations and paving the way for future research to bridge this gap. We release the data and code at https://github.com/claws-lab/persona-modality .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。