arXiv:2607.19341cs.CV2026-07被引 1

构建专家级视觉生成推理基准,评估模型深度知识运用能力。

ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis

论文配图:ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis
图 1 · 摘自论文原文
  • 设计9类认知能力+8个专业领域,形成58个细分任务
  • 推出含1611条专家标注的基准数据集,支持复杂图像生成
  • 提出新训练方法提升模型推理与指令优化,适合研究视觉生成的学者

近年来多模态生成模型使基于指令的图像生成从语义操作迈向知识驱动的视觉推理。然而现有方法仅关注常识推理、浅层因果理解与直接知识回忆,难以应对知识密集型生成任务。为此,我们构建了以能力为中心的基准体系ExpertVerse,涵盖9种认知能力与8个专家领域,共形成58个子领域。我们收集了1,611个由专家标注的实例,覆盖单图编辑、多图合成与文本到图像生成任务。进一步开发自动化流程,生成规模达10万的ExpertVerse-100K数据集,包含推理轨迹与知识锚定的解释标注。基于此,我们训练了采用强化学习微调的KnowThinker模型,具备世界知识的视觉语言推理引擎,可联合生成思考过程与优化指令。针对跨模态奖励错位与多目标梯度冲突问题,提出面向任务的自举帕累托策略优化(BPPO),结合奖励修正与冲突感知优势融合。大量开源与闭源模型实验揭示关键推理缺陷,凸显构建知识密集型评测基准对下一代视觉生成的重要性。

原文摘要 · Abstract (English)

Recent advances in multimodal generative models have enabled instruction-based image generation to move beyond semantic manipulation to knowledge-driven visual reasoning. However, these methods focus on explicit commonsense reasoning, shallow causal understanding, and direct knowledge recall, failing at knowledge-intensive generation. We develop \textbf{ExpertVerse}, a capability-centric benchmark to evaluate generative models via knowledge-intensive lens. ExpertVerse stratifies reasoning generation across an orthogonal taxonomy of \textit{9 cognitive capabilities} and \textit{8 expert disciplines}, yielding \textit{58 sub-disciplines}. We curate 1,611 expert-annotated instances covering single-image editing, multi-image composition, and text-to-image generation. We further develop an automated workflow to produce \textbf{ExpertVerse-100K}, a large-scale dataset with reasoning traces and knowledge-anchored rationale annotations. Based on this, we train \textbf{KnowThinker} with RL fine-tuning, a VLM reasoning engine with world knowledge that jointly generates thinking processes and refined instructions. Towards the cross-modal credit misalignment and multi-objective gradient conflicts in multi-reward optimization, we propose a tailored Bootstrapped Pareto Policy Optimization (BPPO), which synergizes Bootstrapping Reward Rectification (BRR) and Conflict-Aware Pareto Advantage Fusion (CPAF). Extensive results of both open-source and proprietary models exposes critical reasoning deficits, highlighting imperative for knowledge-intensive benchmarks towards next-generation visual generation.

视觉生成知识推理多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。