让故事角色持续更新,不破坏已有角色形象。
EvoTale: Continual Character Customization for Expanding Story Worlds
- 用共享基底的稀疏重叠空间存角色特征,防止互相干扰。
- 根据质量反馈动态调整优化预算,提升定制效率。
- 区域聚焦去噪结合全局保护,兼顾角色身份与剧情连贯。
以角色为中心的故事可视化旨在生成连贯的图像序列,展现叙事事件与人物互动,同时保持角色身份的一致性。在不断扩展的故事世界中,需持续引入用户指定的新角色,尽管存在定制难度差异和多角色场景中的身份冲突,且不能破坏已学习的角色形象。本文提出EvoTale,一种用于扩展风格化故事世界的持续角色定制框架。首先,设计了一种全场景角色集成器(All-in-One-World Character Integrator),通过共享正交基张成的稀疏重叠子空间,在统一的LoRA分支中积累角色特异性残差组件,同时冻结先前学习的组件以限制跨角色耦合。其次,开发了角色质量门(Character Quality Gate),利用基于规则的多模态大模型反馈作为有界控制器,根据评估的定制质量动态调整优化预算。最后,提出角色感知区域聚焦采样(Character-Aware Region-Focus Sampling),结合边界框引导的局部去噪与身份感知的全局去噪,确保角色在其指定区域内保持身份一致,同时维持整体叙事连贯性。实验表明,相比代表性故事可视化与定制方法,EvoTale在角色保真度、持续身份保留、多角色生成质量和故事文本对齐方面取得良好平衡。
原文摘要 · Abstract (English)
Character-centric story visualization aims to synthesize coherent image sequences that depict narrative events and interactions while preserving recurring character identities. In expanding story worlds, new user-specified characters must be continually incorporated despite varying customization difficulty and identity conflicts in multi-character scenes, without disrupting previously learned identities. In this paper, we propose EvoTale, a continual character customization framework for expanding stylized story worlds. We first introduce an All-in-One-World Character Integrator, which accumulates character-specific residual components within a unified LoRA branch using sparsely overlapping subspaces spanned by a shared orthonormal basis, while freezing previously learned components to limit cross-character coupling. We then develop a Character Quality Gate that uses rubric-guided MLLM feedback as a bounded controller to adapt the optimization budget based on the assessed customization quality. Finally, we propose Character-Aware Region-Focus Sampling, which combines bounding-box-guided regional denoising with identity-aware global denoising to preserve character identities within their designated regions while maintaining global narrative coherence. Experimental results show that EvoTale achieves a favorable balance across character fidelity, continual identity retention, multi-character generation quality, and story-text alignment compared with representative story visualization and customization methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。