解决多人视频生成中身份混淆问题,提升角色一致性与互动自然度。
ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
- 用分层注意力机制逐级聚合个体、跨主体及模态特征
- 引入预训练视觉语言模型增强语义理解与关系捕捉能力
- 在线强化学习优化关键概念,适合需要高保真多角色视频的应用
高质量视频合成虽取得显著进展,但现有方法在整合多个角色的身份信息方面仍存在不足,导致语义冲突与身份保留不佳,限制了可控性与实用性。为此,我们提出ID-Crafter框架,实现更优的身份保持与语义连贯性。该框架包含三个核心组件:(i) 分层身份保持注意力机制,逐步融合个体内部、跨主体及跨模态特征;(ii) 基于预训练视觉语言模型(VLM)的语义理解模块,提供细粒度引导并捕捉复杂的跨主体关系;(iii) 在线强化学习阶段,进一步优化模型对关键概念的建模。此外,我们构建了一个新数据集以支持稳健训练与评估。大量实验表明,ID-Crafter在多主体视频生成基准上达到新最优性能,在身份保留、时间一致性与整体视频质量方面均表现优异。
原文摘要 · Abstract (English)
Significant progress has been achieved in high-fidelity video synthesis, yet current paradigms often fall short in effectively integrating identity information from multiple subjects. This leads to semantic conflicts and suboptimal performance in preserving identities and interactions, limiting controllability and applicability. To tackle this issue, we introduce ID-Crafter, a framework for multi-subject video generation that achieves superior identity preservation and semantic coherence. ID-Crafter integrates three key components: (i) a hierarchical identity-preserving attention mechanism that progressively aggregates features at intra-subject, inter-subject, and cross-modal levels; (ii) a semantic understanding module powered by a pretrained Vision-Language Model (VLM) to provide fine-grained guidance and capture complex inter-subject relationships; and (iii) an online reinforcement learning phase to further refine the model for critical concepts. Furthermore, we construct a new dataset to facilitate robust training and evaluation. Extensive experiments demonstrate that ID-Crafter establishes new state-of-the-art performance on multi-subject video generation benchmarks, excelling in identity preservation, temporal consistency, and overall video quality. Project page: https://angericky.github.io/ID-Crafter
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。