arXiv:2512.10943cs.CVcs.AI2025-12被引 4

让视频中多个主体的出现消失时间精确可控,提升个性化视频生成的灵活性。

AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generation

论文配图:AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generation
图 1 · 摘自论文原文
  • 通过显式时间戳编码实现对主体出场时机的精细控制。
  • 在多主体视频生成中首次实现高精度时序控制,视觉质量媲美顶尖方法。
  • 无需额外注意力模块,参数开销极小,适合实际部署。

基于大型扩散模型的主体驱动视频生成技术已能根据用户提供的主体生成个性化内容。然而,现有方法缺乏对主体外观出现与消失的细粒度时间控制,而这对于组合视频合成、故事板制作和可控动画等应用至关重要。我们提出 AlcheMinT,一个统一框架,引入显式时间戳条件以实现主体驱动视频生成。该方法设计了一种新型位置编码机制,可编码时间区间(在此情况下对应主体身份),并无缝融入预训练视频生成模型的位置嵌入。此外,通过引入主体描述性文本标记,强化了视觉身份与视频标题间的绑定关系,缓解生成过程中的歧义问题。通过词级拼接,AlcheMinT 避免了额外的交叉注意力模块,参数开销可忽略不计。我们建立了一个基准,评估多种主体身份保持、视频保真度和时间一致性。实验结果表明,AlcheMinT 在视觉质量上达到当前顶尖视频个性化方法水平,同时首次实现了对多主体视频生成中时序控制的精准把握。

原文摘要 · Abstract (English)

Recent advances in subject-driven video generation with large diffusion models have enabled personalized content synthesis conditioned on user-provided subjects. However, existing methods lack fine-grained temporal control over subject appearance and disappearance, which are essential for applications such as compositional video synthesis, storyboarding, and controllable animation. We propose AlcheMinT, a unified framework that introduces explicit timestamps conditioning for subject-driven video generation. Our approach introduces a novel positional encoding mechanism that unlocks the encoding of temporal intervals, associated in our case with subject identities, while seamlessly integrating with the pretrained video generation model positional embeddings. Additionally, we incorporate subject-descriptive text tokens to strengthen binding between visual identity and video captions, mitigating ambiguity during generation. Through token-wise concatenation, AlcheMinT avoids any additional cross-attention modules and incurs negligible parameter overhead. We establish a benchmark evaluating multiple subject identity preservation, video fidelity, and temporal adherence. Experimental results demonstrate that AlcheMinT achieves visual quality matching state-of-the-art video personalization methods, while, for the first time, enabling precise temporal control over multi-subject generation within videos. Project page is at https://snap-research.github.io/Video-AlcheMinT

视频生成时间控制扩散模型多主体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。