arXiv:2603.25823cs.CVcs.AI2026-03被引 2

测试视觉生成模型的零样本推理能力,发现顶尖模型仍有严重逻辑短板。

ViGoR-Bench: How Far Are Visual Generative Models From Zero-Shot Visual Reasoners?

  • 构建跨模态统一评测框架,覆盖图像到视频任务
  • 同时评估生成过程与最终结果,揭示深层缺陷
  • 用证据驱动的自动评分,更贴近人类判断

现代AIGC模型虽视觉效果惊艳,但在物理、因果或复杂空间推理任务上表现不佳,形成‘逻辑荒漠’。现有评估多依赖表面指标或碎片化基准,导致‘性能幻觉’。为此,我们提出ViGoR(Vision-Generative Reasoning-centric Benchmark)——一个统一评测框架,通过四大创新:1)贯通图像到图像、视频的跨模态覆盖;2)双轨机制评估中间过程与最终输出;3)基于证据的自动化评判,确保高人类对齐;4)细粒度诊断分析,拆解认知维度表现。在20多个领先模型上的实验表明,即使顶尖系统仍存在显著推理缺陷,验证了ViGoR作为下一代智能视觉模型关键‘压力测试’的价值。演示已开放:https://vincenthancoder.github.io/ViGoR-Bench/

原文摘要 · Abstract (English)

Beneath the stunning visual fidelity of modern AIGC models lies a "logical desert", where systems fail tasks that require physical, causal, or complex spatial reasoning. Current evaluations largely rely on superficial metrics or fragmented benchmarks, creating a ``performance mirage'' that overlooks the generative process. To address this, we introduce ViGoR Vision-G}nerative Reasoning-centric Benchmark), a unified framework designed to dismantle this mirage. ViGoR distinguishes itself through four key innovations: 1) holistic cross-modal coverage bridging Image-to-Image and Video tasks; 2) a dual-track mechanism evaluating both intermediate processes and final results; 3) an evidence-grounded automated judge ensuring high human alignment; and 4) granular diagnostic analysis that decomposes performance into fine-grained cognitive dimensions. Experiments on over 20 leading models reveal that even state-of-the-art systems harbor significant reasoning deficits, establishing ViGoR as a critical ``stress test'' for the next generation of intelligent vision models. The demo have been available at https://vincenthancoder.github.io/ViGoR-Bench/

视觉生成推理评测AI评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。