一模型搞定单人双人身份替换,支持任意背景修改
Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation

- 用驱动视频+参考图+文本指令联合控制生成
- 支持多主体替换且无需对齐姿态布局
- 可生成分钟级长视频,适合通用编辑系统
视频身份替换旨在迁移一个或多个主体的身份,同时保持驱动视频的运动、表情和时间结构。现有方法多针对单人场景,常需特定结构控制(如掩码或姿态表示),限制了在通用多模态编辑系统中的灵活性。多人替换进展受限于成对训练数据稀缺。我们提出Vorch-IR,统一框架支持单人与双人身份替换,可选背景替换,仅用一个模型完成。基于LTX2,模型联合条件于驱动视频、索引参考图像及文本编辑指令。参考图像无需匹配驱动视频的姿态、布局或空间配置,其作为主体或背景的用途由指令指定。密集视觉条件通过自注意力融合,视觉-语言上下文通过交叉注意力建立语义对应。我们还开发了自动数据构建流水线,为所有四种编辑场景生成成对监督信号。实验表明,自动指标与成对人工评估均显示强身份保真度、动作保真度和时间连贯性。额外采用时间重叠推理策略,使短片段模型无需自回归续写即可生成分钟级视频。
原文摘要 · Abstract (English)
Video identity replacement seeks to transfer the identities of one or more subjects while preserving the motion, expressions, and temporal structure of a driving video. Existing methods largely target single-person settings and often require task-specific structural controls, such as masks or pose representations, limiting their flexibility in general multimodal editing systems. Progress on multi-person replacement is further constrained by the scarcity of paired training data. We present Vorch-IR, a unified framework that supports single- and dual-person identity replacement, with optional background replacement, in a single model. Built on LTX2, Vorch-IR jointly conditions on a driving video, indexed reference images, and a textual editing instruction. The reference images need not match the pose, layout, or spatial configuration of the driving video: their roles as subject or background references are specified through the instruction. Dense visual conditions are fused through self-attention, while a vision-language context establishes semantic correspondence through cross-attention. We further develop an automatic data construction pipeline that synthesizes paired supervision for all four editing settings. Experiments using automatic metrics and pairwise human evaluation demonstrate strong identity preservation, motion fidelity, and temporal coherence across diverse scenarios. A temporal overlapping inference strategy additionally extends the short-clip model to minute-long generation without autoregressive continuation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。