arXiv:2606.24256cs.CV2026-06

新基准测试揭示视觉模型对罕见物理交互的泛化缺陷

Trimming the Long-Tail of Visual World Modeling Evaluation

论文配图:Trimming the Long-Tail of Visual World Modeling Evaluation
图 1 · 摘自论文原文
  • 构建三类场景评估模型对非常规物理交互的推理能力
  • 模型在异常和不可能场景中性能显著下降,表现严重退化
  • 适合关注世界模型泛化性与真实物理理解的研究者

物理交互呈现长尾分布:少数常见且规律的交互主导人类经验与视觉数据,而大量罕见且不规则的交互却严重缺失。尽管当前视觉世界模型(包括图像与视频生成模型)在现有基准上已实现惊人逼真度,但主要聚焦于模拟常见物理交互。这引出核心问题:当前视觉世界模型是否真正内化并泛化物理原理?本文提出Tailor-Bench基准,挑战模型模拟不规则物理交互的能力。为实现系统评估,设计三种渐进式场景模式:常规场景反映常见工具-任务配对,非常规场景用属性兼容替代品替换常规工具以检验可及性泛化,不可能场景引入属性冲突工具以探测约束意识。此外,在统一评估协议下设计两种互补设置:预测生成要求无引导推断结果,描述生成则指定目标结果以实现忠实再现。实验结果揭示明显的长尾差距:模型性能从常规到非常规再到不可能场景持续下降,表明其泛化能力局限于常见交互。失败分析进一步显示,模型依赖表面视觉模式:图像模型无法正确实现状态变化,视频模型还存在时间不一致性问题。

原文摘要 · Abstract (English)

Physical interactions follow a long-tailed distribution: a set of common and regular interactions dominates human experience and visual data, while a broad spectrum of rare and irregular interactions remains underrepresented. Although recent visual world models, including image and video generation models, achieve impressive realism on existing benchmarks, they primarily focus on simulating common physical interactions. This raises a central question: Do current visual world models internalize and generalize physical principles? In this work, we introduce Tailor-Bench, a benchmark that challenges world models to simulate irregular physical interactions. To enable systematic evaluation, we design three scenario modes that progressively challenge model reasoning: Regular scenarios reflect common tool-task pairs, Unconventional scenarios replace conventional tools with attribute-compatible substitutes to test affordance generalization, and Impossible scenarios introduce attribute-violating tools to probe constraint awareness. Additionally, we design two complementary settings under a unified evaluation protocol: predictive generation requires inferring outcomes without guidance, while descriptive generation specifies the target outcome for faithful realization. Our experimental results reveal a clear long-tail gap in physical world modeling: performance degrades from Regular to Unconventional and Impossible scenarios, indicating limited generalization beyond common interactions. Failure analysis further shows that models rely on superficial visual patterns: image models fail to realize correct state changes, while video models further suffer from temporal inconsistencies.

视觉建模物理理解长尾分布基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。