用强化学习提升多人视频生成的身份一致性。
Identity-GRPO: Optimizing Multi-Human Identity-preserving Video Generation via Reinforcement Learning
- 基于人类反馈构建奖励模型,优化多角色身份保持。
- 在基准上实现最高18.9%的身份一致性提升。
- 适合关注个性化视频生成与角色连贯性的研究者。
尽管VACE和Phantom等先进方法已在特定主体的多样化场景中推进了视频生成,但在动态交互中多角色身份一致性仍面临挑战。为此,我们提出Identity-GRPO,一种由人类反馈驱动的优化流程,用于改进多角色身份保持的视频生成。首先,我们在大规模偏好数据集上训练视频奖励模型,该数据集包含人工标注和合成扰动数据,采用成对标注聚焦于视频全程的人体一致性。随后,我们使用针对多角色一致性定制的GRPO变体,显著提升VACE和Phantom的表现。通过广泛的消融实验,评估了标注质量与设计选择对策略优化的影响。实验表明,Identity-GRPO在人体一致性指标上相较基线方法最高提升18.9%,为强化学习与个性化视频生成的对齐提供了可操作的洞见。
原文摘要 · Abstract (English)
While advanced methods like VACE and Phantom have advanced video generation for specific subjects in diverse scenarios, they struggle with multi-human identity preservation in dynamic interactions, where consistent identities across multiple characters are critical. To address this, we propose Identity-GRPO, a human feedback-driven optimization pipeline for refining multi-human identity-preserving video generation. First, we construct a video reward model trained on a large-scale preference dataset containing human-annotated and synthetic distortion data, with pairwise annotations focused on maintaining human consistency throughout the video. We then employ a GRPO variant tailored for multi-human consistency, which greatly enhances both VACE and Phantom. Through extensive ablation studies, we evaluate the impact of annotation quality and design choices on policy optimization. Experiments show that Identity-GRPO achieves up to 18.9% improvement in human consistency metrics over baseline methods, offering actionable insights for aligning reinforcement learning with personalized video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。