arXiv:2511.23146cs.CV2025-11被引 2

让视频生成能精准控制物体位置和属性,同时保持整体画面一致。

InstanceV: Instance-Level Video Generation

  • 通过实例感知注意力机制,实现物体在指定位置的精准生成
  • 在多个数据集上显著提升实例可控性与整体视频质量
  • 适合需要精细控制视频中物体布局的研究者或应用

近期文本到视频的扩散模型已能根据文本描述生成高质量视频。然而,大多数现有模型仅依赖文本条件,缺乏对视频生成的细粒度控制能力。为此,我们提出InstanceV,一个支持实例级控制和全局语义一致性的视频生成框架。具体而言,借助提出的实例感知掩码交叉注意力机制,InstanceV充分利用额外的实例级定位信息,在指定空间位置生成属性正确的实例。为提升整体一致性,我们引入参数高效的共享时间自适应提示增强模块,以连接局部实例与全局语义。此外,在训练和推理阶段均采用空间感知无条件引导,缓解小尺寸实例消失问题。最后,我们构建了新基准InstanceBench,融合通用视频质量指标与实例感知指标,实现更全面的评估。大量实验表明,InstanceV不仅在实例级控制上表现卓越,还在通用质量和实例感知指标上超越现有最先进模型。

原文摘要 · Abstract (English)

Recent advances in text-to-video diffusion models have enabled the generation of high-quality videos conditioned on textual descriptions. However, most existing text-to-video models rely solely on textual conditions, lacking general fine-grained controllability over video generation. To address this challenge, we propose InstanceV, a video generation framework that enables i) instance-level control and ii) global semantic consistency. Specifically, with the aid of proposed Instance-aware Masked Cross-Attention mechanism, InstanceV maximizes the utilization of additional instance-level grounding information to generate correctly attributed instances at designated spatial locations. To improve overall consistency, We introduce the Shared Timestep-Adaptive Prompt Enhancement module, which connects local instances with global semantics in a parameter-efficient manner. Furthermore, we incorporate Spatially-Aware Unconditional Guidance during both training and inference to alleviate the disappearance of small instances. Finally, we propose a new benchmark, named InstanceBench, which combines general video quality metrics with instance-aware metrics for more comprehensive evaluation on instance-level video generation. Extensive experiments demonstrate that InstanceV not only achieves remarkable instance-level controllability in video generation, but also outperforms existing state-of-the-art models in both general quality and instance-aware metrics across qualitative and quantitative evaluations.

视频生成扩散模型实例控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。