让视频人物全身形象不变,适应任意视角和动作。
WildActor: Unconstrained Identity-Preserving Video Generation

- 用不对称身份保持注意力+视角自适应采样,动态优化参考条件。
- 在1800万张图像、160万段视频上训练,实现跨视角身份一致。
- 适合影视生成、虚拟主播等需稳定人物形象的场景。
高质量数字演员生成要求在动态镜头、多变视角和复杂动作下保持全身形象一致,现有方法常因聚焦面部而忽略身体一致性,或因姿态锁定产生僵硬复制粘贴效果。我们构建了包含1800万张人体图像和160万段视频的Actor-18M大规模数据集,覆盖任意视角与标准三视图。基于此,提出WildActor框架,支持任意视角条件下的真人视频生成。引入非对称身份保持注意力机制与视角自适应蒙特卡洛采样策略,通过边际效用迭代重加权参考条件,实现均衡流形覆盖。在新提出的Actor-Bench评测中,WildActor在多样镜头构图、大视角变化和大幅动作下均显著优于现有方法,持续保持身体身份一致性。
原文摘要 · Abstract (English)
Production-ready human video generation requires digital actors to maintain strictly consistent full-body identities across dynamic shots, viewpoints and motions, a setting that remains challenging for existing methods. Prior methods often suffer from face-centric behavior that neglects body-level consistency, or produce copy-paste artifacts where subjects appear rigid due to pose locking. We present Actor-18M, a large-scale human video dataset designed to capture identity consistency under unconstrained viewpoints and environments. Actor-18M comprises 1.6M videos with 18M corresponding human images, covering both arbitrary views and canonical three-view representations. Leveraging Actor-18M, we propose WildActor, a framework for any-view conditioned human video generation. We introduce an Asymmetric Identity-Preserving Attention mechanism coupled with a Viewpoint-Adaptive Monte Carlo Sampling strategy that iteratively re-weights reference conditions by marginal utility for balanced manifold coverage. Evaluated on the proposed Actor-Bench, WildActor consistently preserves body identity under diverse shot compositions, large viewpoint transitions, and substantial motions, surpassing existing methods in these challenging settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。