arXiv:2511.19836cs.CV2025-11被引 8

构建4D世界生成模型的统一评估框架,涵盖视觉、物理与时空一致性。

4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models

  • 提出四维评估维度:感知质量、条件对齐、物理真实、时空一致。
  • 支持图文视频多模态输入,用大模型辅助判断提升评估可信度。
  • 适配不同输入模态,实现跨模态一致性评价,适合研究世界生成的团队使用。

世界生成模型正成为下一代多模态智能系统的核心。与传统2D图像生成不同,世界模型需从图像、视频或文本生成具有真实感、动态性与物理一致性的3D/4D世界。这些模型不仅需高保真视觉输出,还需在空间、时间、物理规律和指令控制上保持连贯,适用于虚拟现实、自动驾驶、具身智能与内容创作。然而,现有基准侧重不同评估维度,缺乏对世界真实感的统一衡量。为此,我们提出4DWorldBench,从感知质量、条件-4D对齐、物理真实性和4D一致性四个维度系统评估模型,涵盖图像到3D/4D、视频到4D、文本到3D/4D等任务。创新性引入多模态自适应条件映射,将所有输入条件统一转换为文本空间,并融合LLM-as-judge、MLLM-as-judge与传统网络评估方法。该统一自适应设计提升了对对齐性、物理真实性和跨模态一致性的综合评估能力。初步人类实验表明,自适应工具选择更贴近主观判断。本基准旨在推动客观比较与模型改进,加速从‘视觉生成’迈向‘世界生成’。项目主页:https://yeppp27.github.io/4DWorldBench.github.io/

原文摘要 · Abstract (English)

World Generation Models are emerging as a cornerstone of next-generation multimodal intelligence systems. Unlike traditional 2D visual generation, World Models aim to construct realistic, dynamic, and physically consistent 3D/4D worlds from images, videos, or text. These models not only need to produce high-fidelity visual content but also maintain coherence across space, time, physics, and instruction control, enabling applications in virtual reality, autonomous driving, embodied intelligence, and content creation. However, prior benchmarks emphasize different evaluation dimensions and lack a unified assessment of world-realism capability. To systematically evaluate World Models, we introduce the 4DWorldBench, which measures models across four key dimensions: Perceptual Quality, Condition-4D Alignment, Physical Realism, and 4D Consistency. The benchmark covers tasks such as Image-to-3D/4D, Video-to-4D, Text-to-3D/4D. Beyond these, we innovatively introduce adaptive conditioning across multiple modalities, which not only integrates but also extends traditional evaluation paradigms. To accommodate different modality-conditioned inputs, we map all modality conditions into a unified textual space during evaluation, and further integrate LLM-as-judge, MLLM-as-judge, and traditional network-based methods. This unified and adaptive design enables more comprehensive and consistent evaluation of alignment, physical realism, and cross-modal coherence. Preliminary human studies further demonstrate that our adaptive tool selection achieves closer agreement with subjective human judgments. We hope this benchmark will serve as a foundation for objective comparisons and improvements, accelerating the transition from "visual generation" to "world generation." Our project can be found at https://yeppp27.github.io/4DWorldBench.github.io/.

世界生成多模态评估4D一致性基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。