让视频模型在多镜头创作中灵活切换生成、参考与编辑,保持上下文连贯。
ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing

- 用角色感知的上下文表示和固定预算稀疏路由,实现高效历史管理。
- 在三个基础任务上表现优于专用模型,跨镜头一致性显著提升。
- 支持用户自由组合工作流,单卡实现实时16帧/秒推理。
当前视频模型虽能统一生成、参考条件与编辑功能,但通常作为独立操作处理固定输入。实际创作涉及多镜头流程,需模型在文本生成、参考跟随或源素材编辑间切换,并维持共享历史。本文提出交互式多镜头视频创作(IMVC)设定,引入ContextMaster模型,采用角色感知的上下文表示。为避免去噪步骤中上下文读取成本增长,模型结合可复用的干净上下文状态与固定预算稀疏路由,并使用ConstraintSink保持任务约束可见。针对稀疏上下文访问与少步数推断的双重挑战,提出两阶段特权上下文蒸馏框架:先通过一致性蒸馏将密集教师模型的完整上下文行为迁移,再以分布匹配优化部署过程。在三项基础任务上的实验表明,相较专用基线,任务完成度与跨镜头一致性均获提升。用户研究验证了灵活工作流的可行性,模型在单张GPU上达到16 FPS。
原文摘要 · Abstract (English)
Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs. Practical creation unfolds across multiple shots, requiring one model to generate from text, follow a reference, or edit source footage while maintaining shared history. We formalize this setting as interactive multi-shot video creation (IMVC) and introduce ContextMaster, a unified model with a role-aware context representation for these operations. An interactive model must retain access to an expanding history without allowing the context read cost at each denoising step to grow. ContextMaster combines reusable clean context states with fixed budget sparse context routing and uses ConstraintSink to keep task constraints visible. To address the dual challenges of sparse context access and inference with few denoising steps, we propose a two-stage privileged context distillation framework, which transfers full context behavior from a dense teacher through consistency distillation and then refines deployment rollouts with distribution matching. Experiments on the three primitive tasks demonstrate improved task fulfillment and consistency across shots over specialized baselines. User studies further validate flexibly composed workflows, while the model reaches 16 FPS on a single GPU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。