用一张图+文字生成连贯动作视频,保持人物一致且动作自然。
Proteus-ID: ID-Consistent and Motion-Coherent Video Customization
- 融合图像与文本信息构建统一身份表征,避免模态不平衡。
- 动态调节去噪过程中的身份条件,提升细节还原能力。
- 基于光流自监督优化运动生成,无需额外输入即可更真实。
视频身份定制旨在仅凭一张参考图和文本提示,生成特定主体的逼真、时序连贯视频。该任务面临两大挑战:(1)在匹配描述外观与动作的同时保持身份一致性;(2)生成自然流畅的动作,避免僵硬感。为此,我们提出Proteus-ID,一种基于扩散模型的身份一致且动作连贯的视频定制框架。首先,设计多模态身份融合(MIF)模块,通过Q-Former将视觉与文本线索统一为联合身份表征,为扩散模型提供一致引导,消除模态失衡。其次,提出时间感知身份注入(TAII)机制,动态调节去噪步骤中的身份条件,提升细粒度重建。第三,提出自监督的自适应运动学习(AML),根据光流生成的运动热图重加权训练损失,增强运动真实性,无需额外输入。为支持该任务,我们构建了Proteus-Bench数据集,包含20万条精选视频片段,以及150位来自不同职业与族裔的个体用于评估。大量实验表明,Proteus-ID在身份保留、文本对齐与运动质量上均优于现有方法,树立了视频身份定制新基准。代码与数据已公开于https://grenoble-zhang.github.io/Proteus-ID/。
原文摘要 · Abstract (English)
Video identity customization seeks to synthesize realistic, temporally coherent videos of a specific subject, given a single reference image and a text prompt. This task presents two core challenges: (1) maintaining identity consistency while aligning with the described appearance and actions, and (2) generating natural, fluid motion without unrealistic stiffness. To address these challenges, we introduce Proteus-ID, a novel diffusion-based framework for identity-consistent and motion-coherent video customization. First, we propose a Multimodal Identity Fusion (MIF) module that unifies visual and textual cues into a joint identity representation using a Q-Former, providing coherent guidance to the diffusion model and eliminating modality imbalance. Second, we present a Time-Aware Identity Injection (TAII) mechanism that dynamically modulates identity conditioning across denoising steps, improving fine-detail reconstruction. Third, we propose Adaptive Motion Learning (AML), a self-supervised strategy that reweights the training loss based on optical-flow-derived motion heatmaps, enhancing motion realism without requiring additional inputs. To support this task, we construct Proteus-Bench, a high-quality dataset comprising 200K curated clips for training and 150 individuals from diverse professions and ethnicities for evaluation. Extensive experiments demonstrate that Proteus-ID outperforms prior methods in identity preservation, text alignment, and motion quality, establishing a new benchmark for video identity customization. Codes and data are publicly available at https://grenoble-zhang.github.io/Proteus-ID/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。