测试AI如何主动发现视觉规则,揭示现有模型在推理与实验设计上的短板。
Playing ZendoWorld: Challenging AI Agents on Active Visual Concept Induction

- 构建交互式环境ZendoWorld,让智能体通过试错推断隐藏逻辑规则
- 视觉大模型虽能准确预测标签,却无法真正理解规则,且提出无效实验
- 人类表现也受限于复杂规则,提示需改进主动学习能力
构建了一个名为ZendoWorld的受控交互环境,用于研究智能体在感知复杂输入、形成隐藏模式假设并设计有信息量的实验以验证假设方面的能力。我们评估了多种智能体,涵盖纯视觉语言模型推理、贝叶斯粒子滤波、动态概念发现及神经符号方法。主要发现包括:(1)对观测样本的高标签预测准确率并不等同于恢复底层规则;(2)不同智能体类型在感知与归纳环节存在不同的性能瓶颈;(3)基于视觉语言模型的智能体提出的实验几乎无信息量,未能有效降低假设不确定性。为对比分析,我们收集了人类在该任务上的数据,结果显示在复杂规则下仍存在显著的归纳推理差距。总体而言,ZendoWorld为评估智能体的主动性与推理能力提供了重要平台,并指明了科学发现等领域的改进方向。
原文摘要 · Abstract (English)
A central challenge in building intelligent systems is enabling agents to jointly perceive complex inputs, form hypotheses about hidden patterns, and design informative experiments to test them. To study this problem, we propose ZendoWorld, a controlled interactive environment in which agents must infer a logical rule about visual game observations, acquire information by proposing new scenes, and refine their hypotheses based on feedback from the game environment. We evaluate several agents spanning pure VLM reasoning, Bayesian particle filtering, dynamic concept discovery, and neuro-symbolic methods. Our main findings are: (1) high accuracy in predicting labels for observed examples does not imply recovery of the underlying rule; (2) perception and induction are distinct bottlenecks for different agent classes; and (3) VLM-based agents propose near-uninformative experiments, failing to actively reduce hypothesis uncertainty. To compare these results, we collect human data on the task, which reveals a gap in inductive reasoning, particularly for more complex rules. Overall, ZENDOWORLD takes an important step toward evaluating intelligent agents and identifies concrete avenues for improvement, particularly in domains like scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。