多视角视频生成中保持身份一致性的新框架,提升大视角变化下的外观保真度。
HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation

- 通过多层级特征注入,融合正脸参考图与文本提示,稳定保留外观细节。
- 在100个案例、52个身份的多视角数据集上,性能超越开源模型并媲美闭源引擎。
- 适合需要高保真身份一致性视频生成的研究者和开发者。
当前的身份一致视频生成方法在大幅视角变化下难以保持外观保真度。尽管引入多视角参考输入是自然解决方案,但进展受限于缺乏有效框架和多视角数据稀缺。本文提出HarmoView,一种鲁棒的框架,通过三项架构改进与分阶段训练策略,有效整合多视角线索。首先,多层级特征注入(MFI)将正面参考图的ViT原始特征与文本标记通过交叉注意力注入,提供持续的低层外观锚点,增强DiT块中的高层身份特征;其次,可学习代理令牌统一单/多视角参考布局,解决视角不匹配问题;再者,开发跳变旋转位置编码(Jump-RoPE)实现身份特征隔离,减少身份混淆。为激活这些结构能力并保留原有生成先验,提出渐进式视角训练课程(Progressive View Curriculum),采用视图丢弃策略,平稳过渡从标准文本到视频生成到高保真、身份持久的空间推理。此外,构建大规模多视角数据集以缓解数据稀缺问题。在包含100个手动标注案例、52个唯一身份的多视角基准上,大量实验表明HarmoView显著优于开源基线,并达到领先闭源引擎水平,实现身份一致视频生成的最先进性能。
原文摘要 · Abstract (English)
Current identity-consistent video generation methods struggle to preserve appearance fidelity under large viewpoint changes. While introducing multi-view reference input offers a natural solution, progress remains constrained by the lack of effective frameworks for multi-view inputs and the scarcity of multi-view data. We address these challenges by proposing HarmoView, a robust framework for identity-consistent video generation that effectively integrates multi-view cues through three architectural refinements complemented by a staged training curriculum. Specifically, we first introduce Multi-level Feature Injection to anchor identity fidelity; by injecting raw ViT features from frontal references alongside text tokens via cross-attention, MFI provides persistent low-level appearance anchors that complement the high-level identity features within DiT blocks, leading to enhanced identity preservation. Then, we employ learnable proxy tokens to unify heterogeneous reference layouts across single-/multi-view settings while simultaneously resolving the reference-view mismatch problem. Jump-RoPE is further developed for identity-wise feature isolation to reduce identity crosstalk. To activate these structural capabilities while preserving the original generative priors, we propose the Progressive View Curriculum. This four-stage training strategy employs view dropout to facilitate a stable transition from vanilla T2V generation to high-fidelity, identity-persistent spatial reasoning. Furthermore, we construct a large-scale multi-view dataset to address the issue of data scarcity. Extensive evaluation on our multi-view benchmark, comprising 100 manually-curated cases spanning 52 unique identities, demonstrates that HarmoView significantly outperforms open-source baselines and matches leading closed-source engines, achieving state-of-the-art performance in identity-consistent video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。