arXiv:2410.03181cs.CL2024-10EMNLP被引 3

让多模态大模型根据图像人格表现不同谈判行为。

Kiss up, Kick down: Exploring Behavioral Changes in Multi-modal Large Language Models with Assigned Visual Personas

  • 用5000张虚构头像构建视觉人格数据集
  • 模型对攻击性图像反应更激烈,且受对手形象影响
  • 首次揭示视觉人格如何引导模型行为变化

本研究首次探索多模态大语言模型(LLMs)是否能根据视觉人格调整行为,填补了以往文献多关注文本人格的空白。我们构建了一个包含5000张虚构头像的全新数据集,用于赋予模型视觉人格,并分析其在谈判中的行为表现,重点关注攻击性特征。结果表明,模型对图像攻击性的判断与人类相似,当被赋予攻击性视觉人格时,输出更激进的谈判行为。有趣的是,当对手图像显得比自身更不具攻击性时,模型表现出更强攻击性;反之则更克制。

原文摘要 · Abstract (English)

This study is the first to explore whether multi-modal large language models (LLMs) can align their behaviors with visual personas, addressing a significant gap in the literature that predominantly focuses on text-based personas. We developed a novel dataset of 5K fictional avatar images for assignment as visual personas to LLMs, and analyzed their negotiation behaviors based on the visual traits depicted in these images, with a particular focus on aggressiveness. The results indicate that LLMs assess the aggressiveness of images in a manner similar to humans and output more aggressive negotiation behaviors when prompted with an aggressive visual persona. Interestingly, the LLM exhibited more aggressive negotiation behaviors when the opponent's image appeared less aggressive than their own, and less aggressive behaviors when the opponents image appeared more aggressive.

多模态模型人格建模行为分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。