arXiv:2506.22075cs.CV2025-06被引 1

让机器像人一样分快慢思考,用推理时间提升视觉任务表现

Reasoning in machine vision by learning fast and slow thinking

  • 模仿人类双系统认知,快思考生成方案,慢思考迭代优化
  • 推理时间越长,性能越强,超越大模型和人类专家
  • 适合数据少的视觉任务,如癌症定位等医疗场景

推理是人类智能的核心,使我们在复杂陌生场景中灵活决策。相比之下,机器智能仍受限于训练数据,无法在推理时动态优化解决方案。尽管近期研究探索了机器推理——通过增加推理计算量来提升性能——但主要集中在数学等有明确规则的语义领域。许多任务缺乏足够标注数据,需依赖推理时计算来提升性能。本文提出一种视觉领域的机器推理范式,能在有限标注数据下,通过增加推理时间(即推理时计算)持续提升性能。该方法受人类双重认知理论启发,融合快速思考系统(System I)用于熟悉任务的方案生成与验证,以及慢速思考系统(System II)通过自对弈强化学习迭代优化预测,即使缺乏特定任务数据也能有效工作。该范式实现方案提出、竞争与精炼直至收敛。实验表明,在计算机视觉基准测试及跨五类器官的癌症定位任务中,更长的推理时间带来优于大规模监督学习、基础模型和人类专家的性能,凸显推理时计算在数据稀缺问题中的潜力。

原文摘要 · Abstract (English)

Reasoning is a hallmark of human intelligence, enabling adaptive decision-making in complex unfamiliar scenarios. In contrast, machine intelligence remains bound to training data, unable to dynamically refine solutions at inference. While recent advances have explored machine reasoning - trading inference-time compute for improved performance - they focus on verbal domains such as mathematical problem-solving where explicit rules govern step-by-step solution generation. Many tasks lack sufficient labelled data and require alternative performance improvement mechanisms, such as inference-time compute. Here we present a paradigm for machine reasoning in vision, enabling performance improvements with increasing thinking time (inference-time compute), even with limited labelled data. Our approach is inspired by dual-process theories of human cognition, integrating a fast-thinking System I module for generating and verifying solutions in familiar tasks, with a slow-thinking System II module that iteratively refines predictions using self-play reinforcement learning, even when task-specific data is limited. This paradigm involves proposing, competing over, and refining solutions until convergence. We demonstrate that extended inference-time compute yields superior performance compared to large-scale supervised learning, foundation models, and human experts in vision tasks. These include computer-vision benchmarks and cancer localisation across five organs, highlighting the potential of inference-time compute for data-scarce problems.

机器推理视觉任务双系统推理时间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。