arXiv:2602.11144cs.LGcs.AI2026-02被引 13

评测大模型在新情境下的即时推理与创造能力,发现现有模型短板。

GENIUS: Generative Fluid Intelligence Evaluation Suite

  • 构建三项动态能力评估任务:推断隐含模式、执行临时约束、适应上下文知识。
  • 12个主流模型在新场景下表现差,尤其在理解上下文上存在明显缺陷。
  • 提出无需训练的注意力干预方法,可有效提升模型应对复杂情境的能力。

统一多模态模型在视觉生成方面取得显著进展,但现有评测主要聚焦于‘晶体智力’——依赖已有知识和学习范式。这忽视了‘生成式流体智力(GFI)’:即在即时情境中归纳模式、推理约束并灵活适应的能力。为此,我们提出GENIUS(生成式流体智力评估套件),将GFI形式化为三个核心能力:推断隐含模式(如推断个性化视觉偏好)、执行临时约束(如可视化抽象隐喻)、适应上下文知识(如模拟反直觉物理)。这些任务要求模型完全基于当前语境解题。对12个代表性模型的系统评估显示其在这些任务上存在显著性能缺陷。诊断分析表明,问题根源在于上下文理解不足,而非生成能力有限。为此,我们提出一种无需训练的注意力干预策略。GENIUS建立了一套严格的GFI评测标准,推动领域从知识利用迈向动态通用推理。数据集与代码将公开于:https://github.com/arctanxarc/GENIUS。

原文摘要 · Abstract (English)

Unified Multimodal Models (UMMs) have shown remarkable progress in visual generation. Yet, existing benchmarks predominantly assess $\textit{Crystallized Intelligence}$, which relies on recalling accumulated knowledge and learned schemas. This focus overlooks $\textit{Generative Fluid Intelligence (GFI)}$: the capacity to induce patterns, reason through constraints, and adapt to novel scenarios on the fly. To rigorously assess this capability, we introduce $\textbf{GENIUS}$ ($\textbf{GEN}$ Fluid $\textbf{I}$ntelligence Eval$\textbf{U}$ation $\textbf{S}$uite). We formalize $\textit{GFI}$ as a synthesis of three primitives. These include $\textit{Inducing Implicit Patterns}$ (e.g., inferring personalized visual preferences), $\textit{Executing Ad-hoc Constraints}$ (e.g., visualizing abstract metaphors), and $\textit{Adapting to Contextual Knowledge}$ (e.g., simulating counter-intuitive physics). Collectively, these primitives challenge models to solve problems grounded entirely in the immediate context. Our systematic evaluation of 12 representative models reveals significant performance deficits in these tasks. Crucially, our diagnostic analysis disentangles these failure modes. It demonstrates that deficits stem from limited context comprehension rather than insufficient intrinsic generative capability. To bridge this gap, we propose a training-free attention intervention strategy. Ultimately, $\textbf{GENIUS}$ establishes a rigorous standard for $\textit{GFI}$, guiding the field beyond knowledge utilization toward dynamic, general-purpose reasoning. Our dataset and code will be released at: $\href{https://github.com/arctanxarc/GENIUS}{https://github.com/arctanxarc/GENIUS}$.

流体智力多模态评估生成能力上下文理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。