arXiv:2603.20192cs.CVcs.AI2026-03被引 6

让视频生成精准匹配人物属性,实现多角色一致性控制

LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation

论文配图:LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation
图 1 · 摘自论文原文
  • 用关系注意力机制显式建模人物与属性的关联
  • 在多主体视频生成中实现身份一致性和语义对齐
  • 适合需要精细角色控制的视频生成研究者

扩散模型的进展显著提升了文本到视频生成能力,实现了对前景和背景元素的细粒度控制。然而,跨主体的人物属性对齐仍具挑战性,因现有方法缺乏显式机制保障组内一致性。为此,我们提出LumosX框架,在数据与模型设计上双重突破:数据侧通过独立视频的图文协同采集,结合多模态大语言模型推断并标注主体特定依赖关系;这些提取的关系先验构建了更细粒度结构,增强个性化视频生成的表达控制力,并支持建立综合性基准。模型侧引入关系自注意力与关系交叉注意力,将位置感知嵌入与精细化注意力动态融合,显式刻画主体-属性依赖,强化组内凝聚力并提升不同主体簇间的分离度。在自建基准上的全面评估表明,LumosX在细粒度、身份一致性和语义对齐的多主体个性化视频生成任务中达到当前最优性能。代码与模型已公开于https://jiazheng-xing.github.io/lumosx-home/。

原文摘要 · Abstract (English)

Recent advances in diffusion models have significantly improved text-to-video generation, enabling personalized content creation with fine-grained control over both foreground and background elements. However, precise face-attribute alignment across subjects remains challenging, as existing methods lack explicit mechanisms to ensure intra-group consistency. Addressing this gap requires both explicit modeling strategies and face-attribute-aware data resources. We therefore propose LumosX, a framework that advances both data and model design. On the data side, a tailored collection pipeline orchestrates captions and visual cues from independent videos, while multimodal large language models (MLLMs) infer and assign subject-specific dependencies. These extracted relational priors impose a finer-grained structure that amplifies the expressive control of personalized video generation and enables the construction of a comprehensive benchmark. On the modeling side, Relational Self-Attention and Relational Cross-Attention intertwine position-aware embeddings with refined attention dynamics to inscribe explicit subject-attribute dependencies, enforcing disciplined intra-group cohesion and amplifying the separation between distinct subject clusters. Comprehensive evaluations on our benchmark demonstrate that LumosX achieves state-of-the-art performance in fine-grained, identity-consistent, and semantically aligned personalized multi-subject video generation. Code and models are available at https://jiazheng-xing.github.io/lumosx-home/.

视频生成扩散模型个性控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。