用轻量提示路由实现长视频多镜头稳定生成
Prompt-Adapter Context Routing for Parameter-Efficient Multi-Shot Long Video Extrapolation

- 通过低秩时序适配器和提示令牌动态路由控制生成
- 在6个基准上超越现有方法,保持角色/场景/动作一致性
- 适合需要高效长视频生成且资源有限的场景
我们提出PACR-Video,一种参数高效的多镜头长视频外推框架,可在不微调生成器的情况下保持重复实体、场景结构、视觉风格和因果进展。该框架冻结文本到视频扩散变换器,并通过由学习得到的镜头角色提示令牌驱动的低秩时序适配器进行增强。为保证长时程连贯性,构建递归提示库,存储前序镜头的紧凑实体、位置、动作和风格提示,并根据预测的叙事依赖关系通过适配器门控进行路由。采用局部镜头/全局故事联合优化目标,结合下一镜头重建、跨镜头身份对比与提示稀疏正则化,同时通过适配器组合调度平衡早期镜头视觉一致性与后期事件推进及视角变化。在六个多镜头长视频基准上,PACR-Video 在分布质量、语义对齐、身份一致性、时间平滑性、运动稳定性、过渡连贯性和人类偏好等指标上均优于文本到视频、微调型、记忆增强型、流式和递归上下文基线方法。结果表明,紧凑提示路由与轻量级时序适应足以提供稳定长视频外推的可控能力。
原文摘要 · Abstract (English)
We present PACR-Video, a parameter-efficient framework for multi-shot long video extrapolation that preserves recurring entities, scene structure, visual style, and causal progression without full generator fine-tuning. PACR-Video keeps a text-to-video diffusion transformer frozen and augments it with low-rank temporal adapters conditioned by learned shot-role prompt tokens. To maintain long-horizon coherence, it builds a recursive prompt bank that stores compact entity, location, action, and style prompts from previous shots, then routes them through adapter gates according to predicted narrative dependencies. A Shot-Local/Story-Global tuning objective combines next-shot reconstruction, cross-shot identity contrast, and prompt sparsity regularization, while an adapter composition schedule balances early-shot visual consistency with later-shot event progression and viewpoint change. Across six multi-shot and long-video benchmarks, PACR-Video outperforms text-to-video, tuning-based, memory-augmented, streaming, and recursive-context baselines on distributional quality, semantic alignment, identity consistency, temporal smoothness, motion stability, transition coherence, and human preference. These results show that compact prompt routing and lightweight temporal adaptation provide sufficient controllable capacity for stable long video extrapolation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。