无需训练即可精准编辑多主体视频,解决遮罩纠缠与注意力分散问题。
ASTRA: Let Arbitrary Subjects Transform in Video Editing
- 通过提示引导的多模态对齐,缓解多主体注意力稀释。
- 基于先验的遮罩重定向生成连贯时序掩码,解决边界混淆。
- 可即插即用,适合需要高精度多主体编辑的研究与应用。
现有视频编辑方法在单一主体场景表现优异,但在密集多主体场景中常因注意力稀释和遮罩边界纠缠导致属性泄露与时间不稳定。为此,我们提出ASTRA——一种无需训练的无缝任意主体视频编辑框架。ASTRA在不进行模型微调的前提下,能精确操控多个指定主体,严格保留非目标区域。其核心由两个模块构成:提示引导的多模态对齐模块,生成鲁棒条件以缓解注意力稀释;基于先验的遮罩重定向模块,生成时序连贯的掩码序列以解决边界纠缠。作为通用插件式模块,ASTRA可无缝集成于多种掩码驱动的视频生成器。在新构建的基准数据集MSVBench上的大量实验表明,ASTRA持续优于当前最优方法。代码、模型与数据已公开于https://github.com/XWH-A/ASTRA。
原文摘要 · Abstract (English)
While existing video editing methods excel with single subjects, they struggle in dense, multi-subject scenes, frequently suffering from attention dilution and mask boundary entanglement that cause attribute leakage and temporal instability. To address this, we propose ASTRA, a training-free framework for seamless, arbitrary-subject video editing. Without requiring model fine-tuning, ASTRA precisely manipulates multiple designated subjects while strictly preserving non-target regions. It achieves this via two core components: a prompt-guided multimodal alignment module that generates robust conditions to mitigate attention dilution, and a prior-based mask retargeting module that produces temporally coherent mask sequences to resolve boundary entanglement. Functioning as a versatile plug-and-play module, ASTRA seamlessly integrates with diverse mask-driven video generators. Extensive experiments on our newly constructed benchmark, MSVBench, demonstrate that ASTRA consistently outperforms state-of-the-art methods. Code, models, and data are available at https://github.com/XWH-A/ASTRA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。