arXiv:2507.11761cs.CVcs.AI2025-07ICML

一个统一框架让模型一次训练通吃多种视觉推理任务。

Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual Reasoning

  • 把视觉推理问题转化为预测图像可预测性,用统一生成模型解决
  • 单次多任务训练后,模型能零样本推理新任务
  • 适合想减少重训成本的AI研究者

抽象视觉推理(AVR)使人类能快速发现并泛化抽象规则至新场景。设计具备类人AVR能力的智能系统是人工智能领域的长期目标。近期深度AVR求解器在各类任务中取得显著进展,但通常依赖特定任务的设计或参数。在此范式下,解决新任务往往需重新训练模型,甚至调整架构,增加了成本。本文提出一种新型统一条件生成求解器(UCGS),旨在统一框架下处理多种AVR任务。首先,我们证明一些知名AVR任务可重构成评估问题面板中目标图像可预测性的任务。其次,我们说明在该框架下,仅需训练一个条件生成模型即可解决多种任务。实验表明,经过一次多任务训练,UCGS在多种AVR任务中展现出抽象推理能力。尤其,其具备零样本推理能力,可在测试阶段对未见过的AVR任务进行抽象推理。

原文摘要 · Abstract (English)

Abstract visual reasoning (AVR) enables humans to quickly discover and generalize abstract rules to new scenarios. Designing intelligent systems with human-like AVR abilities has been a long-standing topic in the artificial intelligence community. Deep AVR solvers have recently achieved remarkable success in various AVR tasks. However, they usually use task-specific designs or parameters in different tasks. In such a paradigm, solving new tasks often means retraining the model, and sometimes retuning the model architectures, which increases the cost of solving AVR problems. In contrast to task-specific approaches, this paper proposes a novel Unified Conditional Generative Solver (UCGS), aiming to address multiple AVR tasks in a unified framework. First, we prove that some well-known AVR tasks can be reformulated as the problem of estimating the predictability of target images in problem panels. Then, we illustrate that, under the proposed framework, training one conditional generative model can solve various AVR tasks. The experiments show that with a single round of multi-task training, UCGS demonstrates abstract reasoning ability across various AVR tasks. Especially, UCGS exhibits the ability of zero-shot reasoning, enabling it to perform abstract reasoning on problems from unseen AVR tasks in the testing phase.

视觉推理生成模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。