arXiv:2512.12193cs.CV2025-12被引 9

让视频生成同时保持人物外观和动作一致性,突破现有方法局限。

SMRABooth: Subject and Motion Representation Alignment for Customized Video Generation

  • 用自监督编码器和光流编码器提取人物与动作的物体级表征
  • 通过稀疏LoRA注入实现人物与动作解耦,减少相互干扰
  • 在文本到视频生成中显著提升外观与动作的一致性

定制化视频生成旨在生成忠实保留参考图像中主体外观、同时保持参考视频中时序一致动作的视频。现有方法因缺乏对主体和动作的物体级引导,难以同时保证外观相似性和动作模式一致性。为此,我们提出SMRABooth,利用自监督编码器和光流编码器提供物体级主体与动作表征,并在LoRA微调过程中对齐这些表征。方法包含三个核心阶段:(1) 通过自监督编码器提取主体表征,指导主体对齐,增强整体结构与高层语义一致性;(2) 利用光流编码器提取独立于外观的结构连贯、物体级动作轨迹;(3) 提出主体-动作关联解耦策略,通过时空稀疏LoRA注入,有效降低主体与动作LoRA间的干扰。大量实验表明,SMRABooth在主体与动作定制化方面表现优异,能维持稳定的主体外观与动作模式,验证了其在可控文本到视频生成中的有效性。

原文摘要 · Abstract (English)

Customized video generation aims to produce videos that faithfully preserve the subject's appearance from reference images while maintaining temporally consistent motion from reference videos. Existing methods struggle to ensure both subject appearance similarity and motion pattern consistency due to the lack of object-level guidance for subject and motion. To address this, we propose SMRABooth, which leverages the self-supervised encoder and optical flow encoder to provide object-level subject and motion representations. These representations are aligned with the model during the LoRA fine-tuning process. Our approach is structured in three core stages: (1) We exploit subject representations via a self-supervised encoder to guide subject alignment, enabling the model to capture overall structure of subject and enhance high-level semantic consistency. (2) We utilize motion representations from an optical flow encoder to capture structurally coherent and object-level motion trajectories independent of appearance. (3) We propose a subject-motion association decoupling strategy that applies sparse LoRAs injection across both locations and timing, effectively reducing interference between subject and motion LoRAs. Extensive experiments show that SMRABooth excels in subject and motion customization, maintaining consistent subject appearance and motion patterns, proving its effectiveness in controllable text-to-video generation.

视频生成主体一致性动作对齐LoRA微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。