arXiv:2604.01043cs.CV2026-04

让人物和场景自由组合生成视频,控制更准、灵活性更强。

ONE-SHOT: Compositional Human-Environment Video Synthesis via Spatial-Decoupled Motion Injection and Hybrid Context Integration

  • 分离人体动作与环境信息,通过交叉注意力解耦生成
  • 支持分钟级视频生成,保持人物与场景一致性
  • 无需3D预处理,适合快速实用的视频合成

近年来,视频基础模型(VFMs)推动了以人物为中心的视频生成发展,但对主体与场景进行细粒度且独立的编辑仍是关键挑战。现有方法通过刚性3D几何组合增强环境控制时,往往在精确控制与生成灵活性之间存在明显权衡,且依赖繁重的3D预处理,限制了实际可扩展性。本文提出ONE-SHOT,一种参数高效的组合式人-环境视频生成框架。核心思想是将生成过程分解为解耦信号。我们引入规范空间注入机制,通过交叉注意力实现人体动态与环境线索的解耦。提出动态锚定位置编码(Dynamic-Grounded-RoPE),在无启发式3D对齐条件下建立异构空间域间的空间对应关系。为支持长时序生成,设计混合上下文融合机制,保障分钟级生成中主体与场景的一致性。实验表明,该方法显著优于现有最优模型,在结构控制与创意多样性上表现更优。项目已开源:https://martayang.github.io/ONE-SHOT/

原文摘要 · Abstract (English)

Recent advances in Video Foundation Models (VFMs) have revolutionized human-centric video synthesis, yet fine-grained and independent editing of subjects and scenes remains a critical challenge. Recent attempts to incorporate richer environment control through rigid 3D geometric compositions often encounter a stark trade-off between precise control and generative flexibility. Furthermore, the heavy 3D pre-processing still limits practical scalability. In this paper, we propose ONE-SHOT, a parameter-efficient framework for compositional human-environment video generation. Our key insight is to factorize the generative process into disentangled signals. Specifically, we introduce a canonical-space injection mechanism that decouples human dynamics from environmental cues via cross-attention. We also propose Dynamic-Grounded-RoPE, a novel positional embedding strategy that establishes spatial correspondences between disparate spatial domains without any heuristic 3D alignments. To support long-horizon synthesis, we introduce a Hybrid Context Integration mechanism to maintain subject and scene consistency across minute-level generations. Experiments demonstrate that our method significantly outperforms state-of-the-art methods, offering superior structural control and creative diversity for video synthesis. Our project has been available on: https://martayang.github.io/ONE-SHOT/.

视频生成可控生成解耦建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。