让不同动画角色自然互动,保持风格真实不扭曲。
Character Mixing for Video Generation
- 用跨角色嵌入学习角色身份与行为逻辑。
- 合成共存数据提升角色交互质量,避免风格错乱。
- 适合做跨世界角色生成的创作者与研究者。
想象憨豆先生走进《猫和老鼠》——能否生成不同世界角色自然互动的视频?我们研究文本到视频生成中的跨角色交互问题,核心挑战在于保持角色身份与行为特征的同时实现连贯的跨情境互动。这很困难,因为角色从未共存过,且混合风格常导致风格错乱,使写实角色显得卡通或反之。我们提出框架,采用跨角色嵌入(CCE)从多模态源中学习身份与行为逻辑,并引入跨角色增强(CCA)通过合成共存与混合风格数据丰富训练。两者结合使此前从未共现的角色能自然互动,同时保持风格一致性。在包含10个角色的卡通与真人剧集精选基准上实验显示,身份保留、交互质量与抗风格错乱能力显著提升,开启新的生成叙事可能。更多结果与视频见项目页:https://tingtingliao.github.io/mimix/。
原文摘要 · Abstract (English)
Imagine Mr. Bean stepping into Tom and Jerry--can we generate videos where characters interact naturally across different worlds? We study inter-character interaction in text-to-video generation, where the key challenge is to preserve each character's identity and behaviors while enabling coherent cross-context interaction. This is difficult because characters may never have coexisted and because mixing styles often causes style delusion, where realistic characters appear cartoonish or vice versa. We introduce a framework that tackles these issues with Cross-Character Embedding (CCE), which learns identity and behavioral logic across multimodal sources, and Cross-Character Augmentation (CCA), which enriches training with synthetic co-existence and mixed-style data. Together, these techniques allow natural interactions between previously uncoexistent characters without losing stylistic fidelity. Experiments on a curated benchmark of cartoons and live-action series with 10 characters show clear improvements in identity preservation, interaction quality, and robustness to style delusion, enabling new forms of generative storytelling.Additional results and videos are available on our project page: https://tingtingliao.github.io/mimix/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。