让角色看世界:解决多模态角色扮演中的视觉干扰问题
Through the Lens of Character: Resolving Modality-Role Interference in Multimodal Role-Playing Agent

- 通过角色引导的视觉干预,聚焦与角色相关的视觉信息
- 在多个数据集上显著提升角色一致性,最高提升23.6%
- 无需训练,适合希望增强角色扮演真实感的研究者
多模态大语言模型的发展使角色扮演代理(RPAs)进入具身环境。然而人类视觉具有主观性和身份驱动性,而现有多模态模型提取的是通用、去角色化的特征。在角色扮演中,这种通用视觉噪声掩盖了脆弱的角色特征,导致模态-角色干扰(MRI),使代理难以融合视觉依据与角色一致性。为此,我们提出免训练的角色感知视觉干预(CAVI)框架,使代理能以角色视角感知世界。CAVI系统性地缓解MRI:宏观上,角色引导的标记剪枝(CTP)将视觉感受野限制在角色相关实体;微观上,正交特征调制(OFM)将标记投影到角色上下文子空间以提取对齐事实;解码阶段,模态自适应角色引导(MARS)根据视觉依赖动态优化引导强度。大量实验表明,CAVI有效缓解了MRI,显著提升角色一致性的多模态交互。
原文摘要 · Abstract (English)
The advancement of Multimodal Large Language Models (MLLMs) has expanded Role-Playing Agents (RPAs) into visually grounded environments. However, human vision is inherently subjective and identity-driven, whereas existing MLLMs extract objective, character-agnostic features for general tasks. In RPAs, this generic visual noise overpowers fragile character traits, causing Modality-Role Interference (MRI), where agents struggle to integrate visual grounding and character consistency. To address this, we introduce the training-free Character-Aware Visual Intervention (CAVI) framework, enabling agents to perceive the world through the lens of character. CAVI systematically targets MRI: macroscopically, Character-Guided Token Pruning (CTP) restricts the visual receptive field to role-relevant entities; microscopically, Orthogonal Feature Modulation (OFM) projects tokens onto a character-context subspace to extract aligned facts; and during decoding, Modality-Adaptive Role Steering (MARS) dynamically optimizes steering intensity based on visual reliance. Extensive experiments show CAVI effectively alleviates MRI, significantly enhancing character-consistent multimodal interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。