让AI角色在多模态中保持一致人格与风格,仅用10张图即可定制
Towards Customized Multimodal Role-Play

- 分两阶段训练:先统一对齐,再针对角色优化
- 仅需10张图+交互样例,模型即能生成一致的文本与图像
- 适合开发拟人化对话系统或虚拟角色应用
统一的多模态理解与生成模型推动了更丰富的交互体验。然而,在保持跨模态一致性前提下,同时自定义角色的人格、对话风格和视觉形象仍缺乏探索。为此,我们提出新任务——定制化多模态角色扮演(CMRP)。构建了包含20个角色的RoleScape-20数据集,涵盖人物设定、风格描述、视觉/表情线索及图文互动内容。基于统一模型,设计UniCharacter框架,包含统一监督微调(Unified-SFT)与角色专属组相对策略优化(Character-GRPO)。仅需10张图像及对应交互示例,模型即可学习目标角色,并在生成的文本与图像中体现连贯的人格、风格与视觉特征,耗时约100 GPU小时。在RoleScape-20上的实验表明,该方法显著优于现有方法。消融实验进一步验证了跨模态一致性设计与少样本定制策略的有效性。我们认为,结合统一建模的CMRP为下一代具象化、沉浸式交互智能体奠定了基础。
原文摘要 · Abstract (English)
Unified multimodal understanding and generation models enable richer human-AI interaction. Yet jointly customizing a character's persona, dialogue style, and visual identity while maintaining output consistency across modalities remains largely unexplored. To mitigate this gap, we introduce a new task, Customized Multimodal Role-Play (CMRP). We construct the RoleScape-20 dataset comprising 20 characters, including training and evaluation data that cover persona, stylistic descriptions, visual/expressive cues, and text-image interactions. Building on a unified model, we devise UniCharacter, a two-stage training framework containing Unified Supervised Finetuning (Unified-SFT) and character-specific group relative policy optimization (Character-GRPO). Given only 10 images plus corresponding interaction examples, the model acquires the target character and exhibits coherent persona, style, and visual identity in both generated text and images. This process takes about 100 GPU hours. Experiments on the RoleScape-20 dataset show that the proposed method substantially outperforms prior approaches. Ablation studies further validate the effectiveness of our cross-modal consistency design and few-shot customization strategy. We argue that CMRP, coupled with unified modeling, provides a basis for next-generation characterful and immersive interactive agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。