arXiv:2606.11838cs.CV2026-06

用时空场景图增强视频生成的语义对齐能力

Plan-and-Verify Video Reward Reasoning with Spatio-Temporal Scene Graph Grounding

论文配图:Plan-and-Verify Video Reward Reasoning with Spatio-Temporal Scene Graph Grounding
图 1 · 摘自论文原文
  • 将提示分解为原子命题,逐项验证
  • 基于视频提取结构化场景图支撑判断
  • 提升细粒度时序语义对齐,适合生成评估

文本到视频(T2V)生成的奖励模型在后训练阶段常因细粒度语义对齐失败。我们发现现有基于推理的奖励模型存在两个结构性缺陷:未能系统验证提示中每个条件,且视觉证据在自由形式推理中隐含。为此,我们提出SG-PVR,一种基于时空场景图的计划-验证推理框架。验证计划将提示分解为原子命题,确保每项要求被检查。时空场景图从视频中提取,包含实体、属性和时序关联,作为持续的结构化视觉参考。每个命题均基于视频与场景图进行验证,使判断锚定在明确视觉证据上。SG-PVR在语义对齐任务上表现优异,尤其在细粒度时序语义方面。作为测试时重排序器,进一步提升T2V生成的组合对齐能力。

原文摘要 · Abstract (English)

Reward models for text-to-video (T2V) generation guide post-training but often fail at fine-grained semantic alignment. We trace this to two structural weaknesses in existing reasoning-based reward models: they do not systematically verify every condition described in the prompt, and the visual evidence supporting each judgment remains implicit in their free-form reasoning. We propose SG-PVR, a video reward model that addresses these limitations through plan-and-verify reasoning grounded in spatio-temporal scene graphs. The verification plan decomposes the prompt into atomic claims, ensuring every requirement is checked. The spatio-temporal scene graph, encoding entities, attributes, and temporally-grounded relations, is extracted from the video and maintained as a persistent structured visual reference throughout reasoning. Each claim is verified against both the video and the scene graph, anchoring judgments in explicit visual evidence. SG-PVR achieves strong performance on semantic alignment, including fine-grained temporal semantics. As a test-time reranker, it further enhances compositional alignment in T2V generation.

视频生成奖励模型场景图语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。