测试视频生成器是否真懂物理规律,发现它常错得离谱却看起来自然。
PhysWeep: Does a Video Generator Realize the Physics You Ask For?

- 用像素反推物理参数,不依赖标签,直接检测生成结果与请求的偏差。
- 三个模型在六种条件下均出现固定错误值,且受采样种子影响而非用户请求。
- 揭示了现有评估方法无法发现的系统性偏差,适合研究生成模型可信性者阅读。
图像到视频生成器常被称具备隐式物理世界模型,当前社区仅通过视觉合理性评分来验证,但合理性并不等于物理准确。本文提出PhysWeep,一种无需标签的黑箱审计方法:冻结生成器,从生成帧中恢复实际实现的物理参数,评估生成是否可追踪、实际值与请求值偏差多大,并检验文献提出的两种失败机制。使用确定性模拟器作为正控,确认所有指标均可精确测量。在三个开源生成器上,对六组参数扫描进行测试,发现一个此前未记录的可复现故障:当运动可追踪时,两个模型生成的动态高度一致且拟合良好,但收敛至少数几个由采样种子决定的固定错误值,跨两个模型族、两个物理系统及独立追踪器均成立。该现象既非文献预测的全局回退,也非基于案例的钳制,因回退目标依赖种子,且留一法剔除二者;在可估计范围内,保真度斜率统计上无显著差异。平均种子的表现会掩盖此问题——每个样本均被锁定在错误常数上,正是合理度评分结构性盲区所在。代码与协议已公开。
原文摘要 · Abstract (English)
Image-to-video generators are often credited with absorbing physical dynamics as implicit world models, a claim the community currently checks with plausibility scores that ask whether a clip looks consistent with real-world motion. Plausibility is the wrong test on its own, because a clip can look natural while encoding the wrong value of the governing physical parameter, and no existing benchmark measures this gap directly. PhysWeep closes it with a fixed, label-free audit, treating a frozen generator as a black box, recovering the realized parameter from generated pixels, and reporting how often generation is trackable at all, how far the realized value sits from the requested one, and which, if either, of the literature's two proposed failure mechanisms the data support. A deterministic-simulator positive control confirms every score is exactly checkable. Applied to three open generators across six sweep axes, PhysWeep finds a specific, reproducible, previously undocumented failure. Conditional on producing trackable motion, two of the three generate confident, well-fit dynamics that converge to one of a small number of fixed, wrong values selected by the sampling seed rather than by the request, reproducing across two independent model families, two physical systems, and an independent tracker. It matches neither the prior reversion nor the case-based clamping the literature anticipates, because the reversion target is seed-conditional rather than a single global default, and a leave-one-out selection rule rejects both; the in-range faithfulness slope is statistically indistinguishable from zero wherever a response is estimable at all. A benchmark averaging over seeds would never see this: each sample is confidently locked to a wrong constant, exactly the failure a plausibility score is structurally blind to. We release the protocol, suite, and analysis code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。