用手写数字识别测试智能体构建感知状态的能力
MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents

- 将MNIST转为分步观察任务,限制回看以分离感知能力
- 模型在部分可见时表现显著下降,暴露三大瓶颈
- 适合研究具身智能、记忆机制与主动感知的学者
在部分可观测环境中,智能体需协调主动感知与工作记忆以维持动态的感知状态。然而,现有基准因引入物理和控制复杂性,难以独立评估感知状态的构建与解读能力。为此,我们提出MNIST-PRO,通过将MNIST数字识别转化为带回看约束的分步窥视任务,隔离了智能体感知能力。我们在四种记忆表示(原始视觉历史、文本状态、结构化度量网格地图、整合视觉画布)下评估了十种多模态模型。尽管模型在全可观测下表现良好,但在部分可观测条件下出现明显性能差距。我们识别出三个关键瓶颈:第一,感知状态的构建与解读困难,智能体难以整合碎片化瞥见;第二,智能体常在未完成序列前停止探索;第三,即使面对后续矛盾证据,模型仍无法修正早期错误信念。结果表明,仅获取视觉证据不足,智能体还需具备构建并更新可靠感知状态的能力。
原文摘要 · Abstract (English)
AI agents in partially observable environments need to coordinate active sensing with working memory to maintain an evolving perceptual state. However, existing benchmarks struggle to isolate this perceptual-state construction and interpretation capability because they introduce physical and control complexities. We address this with MNIST-PRO, a benchmark that isolates agentic perception by converting MNIST digit recognition into a sequential, glimpse-based search task with lookback constraints. We evaluate ten multimodal models across four memory representations, including raw visual history, textual states, structured metric grid maps, and a consolidated visual canvas. While models excel under full observability, partial observability exposes a clear performance gap. We identify three distinct bottlenecks. First, perceptual-state construction and interpretation present a challenge, as agents struggle to integrate fragmented glimpses. Second, agents often stop exploring before they see the full sequence. Third, models often fail to revise early, incorrect beliefs even when faced with subsequent contradictory evidence. These results show that simply acquiring visual evidence is not enough. Agents must also be able to build and update a reliable perceptual state.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。