arXiv:2602.22745cs.CV2026-02

让视频生成更准确地遵循文本中的动态空间关系。

SPATIALALIGN: Aligning Dynamic Spatial Relationships in Video Generation

  • 用几何度量法评估生成视频与文本描述的空间关系对齐程度
  • 在多个动态空间关系任务上显著优于基线模型
  • 适合关注视频生成准确性与逻辑一致性的研究者

多数文本到视频生成模型注重视觉美感,却常忽略生成视频中的空间约束。本文提出SPATIALALIGN,一个自提升框架,增强文本到视频模型对文本提示中指定的动态空间关系(DSR)的刻画能力。我们采用零阶正则化直接偏好优化(DPO)微调模型,以更好地对齐动态空间关系。为此,设计了基于几何的DSR-SCORE指标,定量评估生成视频与提示中动态空间关系的匹配程度,突破了以往依赖视觉语言模型(VLM)进行评估的局限。同时构建了一个包含多样化动态空间关系的文本-视频数据集,以支持相关研究。大量实验表明,经微调后的模型在空间关系一致性方面显著优于基线。代码将公开于链接。项目主页:https://fengming001ntu.github.io/SpatialAlign/

原文摘要 · Abstract (English)

Most text-to-video (T2V) generators prioritize aesthetic quality, but often ignoring the spatial constraints in the generated videos. In this work, we present SPATIALALIGN, a self-improvement framework that enhances T2V models capabilities to depict Dynamic Spatial Relationships (DSR) specified in text prompts. We present a zeroth-order regularized Direct Preference Optimization (DPO) to fine-tune T2V models towards better alignment with DSR. Specifically, we design DSR-SCORE, a geometry-based metric that quantitatively measures the alignment between generated videos and the specified DSRs in prompts, which is a step forward from prior works that rely on VLM for evaluation. We also conduct a dataset of text-video pairs with diverse DSRs to facilitate the study. Extensive experiments demonstrate that our fine-tuned model significantly out performs the baseline in spatial relationships. The code will be released in Link. Project page: https://fengming001ntu.github.io/SpatialAlign/

视频生成空间关系文本到视频评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。