arXiv:2511.11134cs.AI2025-11被引 6

构建几何生成推理基准,评估多模态模型的跨模态主动构造能力。

GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models

  • 以几何构图为任务,融合语言理解与精准视觉生成
  • 首次系统评测模型在生成过程中进行推理与构造的能力
  • 适合研究多模态生成、认知推理的学者与工程师

统一多模态模型(UMMs)标志着人工智能范式的转变,从被动感知转向主动的跨模态生成。尽管其信息整合能力空前,但评估仍存在关键空白:现有基准主要分别评估判别性理解或无约束图像生成,无法衡量生成推理的综合认知过程。为此,我们提出几何构图是理想的测试场景,因其天然要求语言理解与精确视觉生成的融合。本文提出GGBench,一个专为评估几何生成推理而设计的基准。它提供了一个系统框架,用于诊断模型不仅理解与推理,还能主动构造解决方案的能力,从而为下一代智能系统设定更严格的标准。项目网站:https://opendatalab-raiser.github.io/GGBench/

原文摘要 · Abstract (English)

The advent of Unified Multimodal Models (UMMs) signals a paradigm shift in artificial intelligence, moving from passive perception to active, cross-modal generation. Despite their unprecedented ability to synthesize information, a critical gap persists in evaluation: existing benchmarks primarily assess discriminative understanding or unconstrained image generation separately, failing to measure the integrated cognitive process of generative reasoning. To bridge this gap, we propose that geometric construction provides an ideal testbed as it inherently demands a fusion of language comprehension and precise visual generation. We introduce GGBench, a benchmark designed specifically to evaluate geometric generative reasoning. It provides a comprehensive framework for systematically diagnosing a model's ability to not only understand and reason but to actively construct a solution, thereby setting a more rigorous standard for the next generation of intelligent systems. Project website: https://opendatalab-raiser.github.io/GGBench/.

多模态生成推理基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。