提出自适应搜索策略,让视频生成更懂天马行空的创意提示。
ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints
- 根据提示语语义关系动态调整搜索空间和奖励函数
- 在2839个长距离语义组合上显著提升生成质量
- 适合研究创意视频生成与测试时优化的学者
视频生成模型在真实场景中表现优异,但在包含稀有共现概念和远距离语义关系的想象性提示下性能明显下降,因这些情况超出训练分布。现有方法多采用固定的测试时缩放,受限于静态搜索空间和奖励设计,难以适应创造性场景。为此,本文提出ImagerySearch,一种基于提示的自适应测试时搜索策略,可动态调整推理搜索空间与奖励函数以匹配提示中的语义关系,从而生成更连贯、视觉合理的想象视频。为评估该方向进展,我们构建了首个针对长距离语义提示的基准LDT-Bench,包含2,839个多样化概念对,并提供自动化评测协议。大量实验表明,ImagerySearch在LDT-Bench上持续优于主流视频生成基线及现有测试时缩放方法,在VBench上也取得竞争力提升,验证了其在多种提示类型下的有效性。代码与数据集将公开,推动想象力驱动视频生成研究。
原文摘要 · Abstract (English)
Video generation models have achieved remarkable progress, particularly excelling in realistic scenarios; however, their performance degrades notably in imaginative scenarios. These prompts often involve rarely co-occurring concepts with long-distance semantic relationships, falling outside training distributions. Existing methods typically apply test-time scaling for improving video quality, but their fixed search spaces and static reward designs limit adaptability to imaginative scenarios. To fill this gap, we propose ImagerySearch, a prompt-guided adaptive test-time search strategy that dynamically adjusts both the inference search space and reward function according to semantic relationships in the prompt. This enables more coherent and visually plausible videos in challenging imaginative settings. To evaluate progress in this direction, we introduce LDT-Bench, the first dedicated benchmark for long-distance semantic prompts, consisting of 2,839 diverse concept pairs and an automated protocol for assessing creative generation capabilities. Extensive experiments show that ImagerySearch consistently outperforms strong video generation baselines and existing test-time scaling approaches on LDT-Bench, and achieves competitive improvements on VBench, demonstrating its effectiveness across diverse prompt types. We will release LDT-Bench and code to facilitate future research on imaginative video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。