arXiv:2512.12375cs.CV2025-12被引 1

无需训练,用少量图片实现视频主角高保真个性化生成。

V-Warper: Appearance-Consistent Video Diffusion Personalization via Value Warping

  • 仅用几张参考图,通过轻量级编码和嵌入适配实现主体身份建模。
  • 推理时利用语义对应关系精准调整视觉特征,提升帧间外观一致性。
  • 适合需要快速个性化视频生成且资源有限的场景。

视频个性化旨在生成能忠实反映用户提供的主体并遵循文本提示的视频。然而,现有方法通常依赖昂贵的视频微调或大规模视频数据集,计算成本高且难以扩展。同时,在帧间保持精细外观一致性仍具挑战。为此,我们提出 V-Warper,一种针对基于变压器的视频扩散模型的免训练粗到精个性化框架。该框架在不进行额外视频训练的情况下增强细粒度身份保真度。(1)轻量级粗略外观适应阶段仅使用任务所需的少量参考图像,通过仅图像的 LoRA 和主体嵌入适配编码全局主体身份。(2)推理时的精细外观注入阶段通过无 RoPE 的中层查询-键特征计算语义对应关系,引导富含外观信息的值表示向生成过程中的语义对齐区域进行变形,掩码确保空间可靠性。V-Warper 显著提升外观保真度,同时保持提示对齐与运动动态,并在无需大规模视频微调的情况下高效实现这些改进。

原文摘要 · Abstract (English)

Video personalization aims to generate videos that faithfully reflect a user-provided subject while following a text prompt. However, existing approaches often rely on heavy video-based finetuning or large-scale video datasets, which impose substantial computational cost and are difficult to scale. Furthermore, they still struggle to maintain fine-grained appearance consistency across frames. To address these limitations, we introduce V-Warper, a training-free coarse-to-fine personalization framework for transformer-based video diffusion models. The framework enhances fine-grained identity fidelity without requiring any additional video training. (1) A lightweight coarse appearance adaptation stage leverages only a small set of reference images, which are already required for the task. This step encodes global subject identity through image-only LoRA and subject-embedding adaptation. (2) A inference-time fine appearance injection stage refines visual fidelity by computing semantic correspondences from RoPE-free mid-layer query--key features. These correspondences guide the warping of appearance-rich value representations into semantically aligned regions of the generation process, with masking ensuring spatial reliability. V-Warper significantly improves appearance fidelity while preserving prompt alignment and motion dynamics, and it achieves these gains efficiently without large-scale video finetuning.

视频生成扩散模型个性化免训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。