arXiv:2412.16495cs.CVcs.MM2024-12被引 8

无需微调,通过姿态引导实现多角色文本生成视频

Follow-Your-MultiPose: Tuning-Free Multi-Character Text-to-Video Generation via Pose Guidance

  • 分离文本与姿态引导,用大模型生成角色专属提示
  • 通过空间对齐注意力和多分支控制模块实现精准多角色生成
  • 支持多种个性化图像生成模型,适用于真实场景多角色视频

文本可编辑且姿态可控的角色视频生成是具有实际应用价值的挑战性课题。然而,现有方法主要聚焦于单角色视频生成,忽略了现实中多个角色同时出现的场景。为此,我们提出一种无需微调的多角色视频生成框架,基于分离的文本与姿态引导机制。首先,从姿态序列中提取角色掩码以确定每个角色的空间位置,再利用大语言模型为每个角色生成独立提示,实现精确的文本引导。此外,提出空间对齐交叉注意力和多分支控制模块,实现细粒度可控的多角色视频生成。可视化结果表明,该方法在多角色生成中具备高精度可控性。我们还验证了其在多种个性化文本到图像模型上的通用性,定量结果表明,本方法优于现有工作。

原文摘要 · Abstract (English)

Text-editable and pose-controllable character video generation is a challenging but prevailing topic with practical applications. However, existing approaches mainly focus on single-object video generation with pose guidance, ignoring the realistic situation that multi-character appear concurrently in a scenario. To tackle this, we propose a novel multi-character video generation framework in a tuning-free manner, which is based on the separated text and pose guidance. Specifically, we first extract character masks from the pose sequence to identify the spatial position for each generating character, and then single prompts for each character are obtained with LLMs for precise text guidance. Moreover, the spatial-aligned cross attention and multi-branch control module are proposed to generate fine grained controllable multi-character video. The visualized results of generating video demonstrate the precise controllability of our method for multi-character generation. We also verify the generality of our method by applying it to various personalized T2I models. Moreover, the quantitative results show that our approach achieves superior performance compared with previous works.

视频生成多角色姿态引导无微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。