用问答图评估视频物理合理性,精准定位违规点。
Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation

- 构建问答图模型,分层检验物体、动作与物理规律一致性。
- 在多个主流模型上验证,与人工评分相关性优于已有方法。
- 适合关注视频生成物理真实性的研究者与开发者。
视频生成模型虽日益逼真,但在遵循基本物理规律方面仍存不足,且缺乏可靠的细粒度评估手段来定位具体违规。为此,我们提出物理问题场景图(PQSG),一种基于层级问答的评估框架。PQSG利用视觉语言模型生成高质量上下文提示下的问题图,通过图结构建模问题间的逻辑依赖,确保查询上下文有效,并评估生成视频在对象、动作及物理规律上的忠实度。我们构建了FinePhyEval数据集,包含多种物理提示和来自Sora 2、Veo 3、Wan 2.1等先进模型生成的视频,由人工标注多维度违规类别。实验表明,PQSG评分与人工判断相关性更高;且其对闭源模型的物理真实感评价高于Wan 2.1。此外,数据集还可用于子任务评估:我们在生成与回答问题上基准测试两个强视觉语言模型,发现模型可生成类人问题,但答问能力仍逊于人类。
原文摘要 · Abstract (English)
Video generation models are increasingly capable of producing realistic videos, but they still struggle to generate videos that follow basic physical laws. Compounding this is a lack of reliable granular evaluation methods for localizing and specifying physical law violations in videos. We address this by introducing Physics Question Scene Graph (PQSG), a hierarchical question-based evaluation pipeline. PQSG evaluates generated videos by checking their faithfulness to a prompt across objects, actions, and adherence to physical laws using a graph-based hierarchy of questions generated by a vision-language model (VLM), guided by high-quality in-context examples. By representing questions as a graph, PQSG introduces logical dependencies within questions, ensuring that each query is contextually valid. Moreover, PQSG provides granular assessments of which qualities of the video violate physical plausibility constraints. We validate PQSG by creating FinePhyEval, a dataset with physics-based prompts and corresponding generated videos from diverse state-of-the-art video generation models (Sora 2, Veo 3, and Wan 2.1), with each video annotated across multiple categories by humans. Using FinePhyEval, we measure the correlation between PQSG's fine-grained scores and human judgments, showing higher overall correlations than prior work. We also find that PQSG ranks closed-source models higher than Wan 2.1 on physical realism. Lastly, we show that the annotations we provide in FinePhyEval can also be used for subtask evaluation: we benchmark two strong VLMs on generating and answering questions, finding that while models can create human-like questions, they still fall short of human performance in answering them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。