arXiv:2606.11074cs.CLcs.AI2026-06

让视觉语言模型学会多性格切换,更好模拟复杂社交行为。

Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models

论文配图:Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models
图 1 · 摘自论文原文
  • 通过显式人格条件控制模型行为,实现性格诱导与切换。
  • 多性格共存时模型表现受前后性格共同影响,存在平衡与残余效应。
  • 适合研究多模态模型人格建模、社交交互与可控生成的学者。

随着多模态大语言模型(MLLMs)在社交交互中的广泛应用,理解并控制其在复杂人格情境下的行为至关重要。本文提出显式人格条件化方法,并建立涵盖单人格诱导、多人格诱导及人格切换的系统性评估框架。实验表明,人格诱导可提升图像描述性能,但会损害需要精确推理的任务表现,如视觉问答(VQA)。在多特质组合与动态切换过程中观察到平衡与残余效应,表明模型行为由先前与当前人格约束共同调制。现有基于提示的人格诱导方法在多模态场景中迁移能力有限。本工作揭示了MLLM中人格建模的动态复杂性,强调需开发更稳健、定制化的诱导与评估方法。代码将在论文接收后公开。

原文摘要 · Abstract (English)

With the widespread deployment of Multimodal Large Language Models (MLLMs) in social interaction, understanding and controlling their behavior under complex personality conditions is essential. This paper introduces explicit personality conditioning and establishes a systematic evaluation framework encompassing single-personality induction, multi-personality induction, and personality switching. Experiments show that personality induction improves image captioning performance but can impair performance on tasks requiring precise reasoning, such as visual question answering (VQA). Balancing and residual effects are observed during multi-trait composition and dynamic switching, indicating that model behavior is co-modulated by both previous and current personality constraints. Existing prompt-based personality induction methods show limited transferability to multimodal settings. Our work reveals the dynamic and complex nature of personality modeling in MLLMs and underscores the need for robust, tailored methods for personality induction and evaluation. The code will be released when the paper is accepted.

人格建模多模态可控生成视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。