arXiv:2512.12756cs.CV2025-12被引 4

首个支持任意模态双向交互的多模态评测基准,覆盖图文音视频全场景。

FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning

  • 构建跨模态互补筛选策略,生成融合依赖的多模态数据。
  • 涵盖16项任务、3268个样本,覆盖40+高质量数据源。
  • 适合评估下一代全模态模型在理解、生成与推理上的综合能力。

尽管多模态大语言模型(MLLMs)和全模态架构迅速发展,现有评测基准仍存在模态覆盖不全、交互以文本为中心、模态间依赖性弱等问题。为此,我们提出FysicsWorld,首个支持图像、视频、音频与文本之间任意双向输入输出的统一全模态评测基准,实现理解、生成与推理的全面评估。该基准包含16项主任务和3268个精选样本,来自40多个高质量数据源,覆盖丰富开放领域类别与多样化题型。我们提出跨模态互补筛选(CMCS)策略,集成于系统化数据构建框架中,生成支持语音交互与融合依赖的跨模态推理数据。对30余种先进基线模型(包括MLLMs、单模态模型、统一理解-生成模型及全模态语言模型)的全面评估揭示了各模型在理解、生成与推理中的性能差异与局限。本基准为下一代全模态架构的评测与推进提供了统一基础与强基线。

原文摘要 · Abstract (English)

Despite rapid progress in multimodal large language models (MLLMs) and emerging omni-modal architectures, current benchmarks remain limited in scope and integration, suffering from incomplete modality coverage, restricted interaction to text-centric outputs, and weak interdependence and complementarity among modalities. To bridge these gaps, we introduce FysicsWorld, the first unified full-modality benchmark that supports bidirectional input-output across image, video, audio, and text, enabling comprehensive any-to-any evaluation across understanding, generation, and reasoning. FysicsWorld encompasses 16 primary tasks and 3,268 curated samples, aggregated from over 40 high-quality sources and covering a rich set of open-domain categories with diverse question types. We also propose the Cross-Modal Complementarity Screening (CMCS) strategy integrated in a systematic data construction framework that produces omni-modal data for spoken interaction and fusion-dependent cross-modal reasoning. Through a comprehensive evaluation of over 30 state-of-the-art baselines, spanning MLLMs, modality-specific models, unified understanding-generation models, and omni-modal language models, FysicsWorld exposes the performance disparities and limitations across models in understanding, generation, and reasoning. Our benchmark establishes a unified foundation and strong baselines for evaluating and advancing next-generation full-modality architectures.

多模态评测基准全模态跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。