arXiv:2606.26058cs.CV2026-06被引 1

让视频生成自由切换领域,保持主体一致且灵活适配文本描述。

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation

论文配图:DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation
图 1 · 摘自论文原文
  • 分离视频与参考特征,用领域感知的AdaLN实现跨域建模。
  • 引入双RoPE机制,精准控制主体空间位置,提升生成一致性。
  • 适用于需要高保真主体和多风格变换的开放领域视频生成任务。

开放领域主体驱动文本到视频生成受到学术界和产业界广泛关注。该任务主要包含两种场景:域内场景需尽可能保留参考主体特征,而跨域场景则在保持主体内在特征的同时,允许主体无关属性根据文本提示灵活变化。现有方法主要聚焦于域内场景下的主体保真度,限制了其在跨域场景中的编辑性和适应性,如新风格、语义组合或领域属性变化。本文提出理想S2V方法应能自由穿梭于不同领域,在域内与跨域场景中均表现优异。为此,我们提出DomainShuttle,可实现开放领域视频个性化的高保真与强灵活性。具体地,我们引入域间解耦的Domain-MoT,将视频与参考特征分离,并采用领域感知的AdaLN进行特定领域建模;提出Video-Reference DualRoPE方案,将参考图像令牌与视频令牌置于独立的RoPE空间,实现主体级别的精确空间建模;设计Cross-Pair Consistent Loss,旨在提取不受无关特征干扰的主体内在特征。大量实验表明,DomainShuttle在多种开放领域应用场景中显著优于现有方法,兼具高主体保真度与强生成灵活性。

原文摘要 · Abstract (English)

Open domain subject-driven text-to-video (S2V) generation has drawn significant interest in academia and industry. Open domain S2V mainly involves two scenarios: in-domain, which requires retaining the reference subject features as much as possible, and cross-domain, which preserves the intrinsic features of the subject while allowing subject-irrelevant properties to vary flexibly according to the text prompt. Existing methods primarily focus on maximizing subject fidelity in in-domain scenarios, which limits their editability and adaptability in cross-domain scenarios, such as novel styles, semantic combinations, or domain attributes. In this study, we propose that an ideal S2V method should flexibly shuttle between different domains, achieving strong performance in both in-domain and cross-domain scenarios. To this end, we propose DomainShuttle, which could achieve high fidelity and generative flexibility for open domain video personalization. Specifically, we introduce Domain-MoT, which decouples videos and reference features and introduces the domain-aware AdaLN for domain-specific modeling of reference images. We then introduce the Video-Reference DualRoPE scheme, which places reference image tokens and video tokens in separate RoPE spaces to enable precise subject-level spatial modeling, and Cross-Pair Consistent Loss, which aims to extract intrinsic subject features unaffected by irrelevant features. Extensive experiments demonstrate that DomainShuttle achieves significant performance improvements over existing methods, exhibiting high subject fidelity and generative flexibility across diverse open domain application scenarios.

视频生成文本生成主体保真跨域建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。