arXiv:2502.11079cs.CVcs.AI2025-02ICCV被引 124

让视频生成保持主体一致性,参考图+文字指令也能精准还原人物特征。

Phantom: Subject-consistent video generation via cross-modal alignment

论文配图:Phantom: Subject-consistent video generation via cross-modal alignment
图 1 · 摘自论文原文
  • 用文本-图像-视频三元组训练跨模态对齐模型,实现双模态融合
  • 生成视频在人物特征上高度一致,避免图像信息泄露和多主体混淆
  • 适合需要角色一致性的人像视频生成,如虚拟主播、数字人

视频生成基础模型持续发展,但主体一致性生成仍处于探索阶段。我们提出Subject-to-Video任务,从参考图像中提取主体特征,并根据文本指令生成保持主体一致的视频。核心在于平衡文本与图像双模态提示,实现文本与视觉内容的深度协同对齐。为此,我们提出Phantom统一框架,支持单主体与多主体参考输入。基于现有文生视频与图生视频架构,重新设计联合文本-图像注入模块,通过文本-图像-视频三元组数据驱动跨模态对齐学习。所提方法在保证高保真度的同时,有效解决图像内容泄露与多主体混淆问题。评估显示,其性能超越其他最先进的闭源商业方案。尤其在人物生成中强调主体一致性,覆盖现有身份保留视频生成能力,并具备更优表现。

原文摘要 · Abstract (English)

The continuous development of foundational models for video generation is evolving into various applications, with subject-consistent video generation still in the exploratory stage. We refer to this as Subject-to-Video, which extracts subject elements from reference images and generates subject-consistent videos following textual instructions. We believe that the essence of subject-to-video lies in balancing the dual-modal prompts of text and image, thereby deeply and simultaneously aligning both text and visual content. To this end, we propose Phantom, a unified video generation framework for both single- and multi-subject references. Building on existing text-to-video and image-to-video architectures, we redesign the joint text-image injection model and drive it to learn cross-modal alignment via text-image-video triplet data. The proposed method achieves high-fidelity subject-consistent video generation while addressing issues of image content leakage and multi-subject confusion. Evaluation results indicate that our method outperforms other state-of-the-art closed-source commercial solutions. In particular, we emphasize subject consistency in human generation, covering existing ID-preserving video generation while offering enhanced advantages.

视频生成主体一致跨模态对齐数字人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。