无需训练,实时优化视频生成的组合能力。
TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
- 引入新参数,通过布局注意力目标优化生成过程。
- 在T2V-CompBench和Vbench上显著提升组合视频生成效果。
- 支持灵活记忆操作,适合需要快速对齐文本的场景。
视频基础模型(VFMs)在视觉生成方面表现优异,但在组合场景(如运动、数量、空间关系)中表现不佳。本文提出测试时优化与记忆(TTOM),一种无需训练的框架,在推理阶段通过优化新参数来对齐视频生成与时空布局,实现更好的文本-图像对齐。不同于现有方法直接干预潜在表示或逐样本调整注意力,我们采用通用布局-注意力目标引导新参数的集成与优化。此外,将视频生成建模为流式任务,利用参数化记忆机制保存历史优化上下文,支持插入、读取、更新、删除等灵活操作。实验表明,TTOM能解耦组合世界知识,展现出强大的迁移性与泛化能力。在T2V-CompBench和Vbench基准上的结果证明,该框架在不依赖训练的情况下,可高效、可扩展地实现组合视频生成的跨模态对齐。
原文摘要 · Abstract (English)
Video Foundation Models (VFMs) exhibit remarkable visual generation performance, but struggle in compositional scenarios (e.g., motion, numeracy, and spatial relation). In this work, we introduce Test-Time Optimization and Memorization (TTOM), a training-free framework that aligns VFM outputs with spatiotemporal layouts during inference for better text-image alignment. Rather than direct intervention to latents or attention per-sample in existing work, we integrate and optimize new parameters guided by a general layout-attention objective. Furthermore, we formulate video generation within a streaming setting, and maintain historical optimization contexts with a parametric memory mechanism that supports flexible operations, such as insert, read, update, and delete. Notably, we found that TTOM disentangles compositional world knowledge, showing powerful transferability and generalization. Experimental results on the T2V-CompBench and Vbench benchmarks establish TTOM as an effective, practical, scalable, and efficient framework to achieve cross-modal alignment for compositional video generation on the fly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。