arXiv:2602.16826cs.LGcs.AI2026-02中稿 · the Workshop on Th…

用分层隐变量模型提升AI在复杂场景中推理他人意图的能力

HiVAE: Hierarchical Latent Variables for Scalable Theory of Mind

  • 构建三层变分自编码器,模仿人类信念-欲望-意图认知结构
  • 在3185个节点的校园导航任务中显著提升心理状态预测性能
  • 提出自监督对齐方法,呼吁社区探讨如何让隐变量真正对应真实心智状态

心智理论(ToM)使人工智能系统能够推断智能体的隐藏目标与心理状态,但现有方法主要局限于小规模、易理解的网格世界。我们提出HiVAE,一种分层变分架构,将心智理论推理扩展至真实的时空域。受人类认知中信念-欲望-意图结构启发,我们的三层变分自编码器在3,185节点的校园导航任务中实现显著性能提升。然而,我们发现一个关键局限:尽管分层结构改善了预测效果,学习到的隐变量缺乏与实际心理状态的显式关联。为此,我们提出自监督对齐策略,并发布此工作以征求社区对对齐方法的反馈。

原文摘要 · Abstract (English)

Theory of mind (ToM) enables AI systems to infer agents' hidden goals and mental states, but existing approaches focus mainly on small human understandable gridworld spaces. We introduce HiVAE, a hierarchical variational architecture that scales ToM reasoning to realistic spatiotemporal domains. Inspired by the belief-desire-intention structure of human cognition, our three-level VAE hierarchy achieves substantial performance improvements on a 3,185-node campus navigation task. However, we identify a critical limitation: while our hierarchical structure improves prediction, learned latent representations lack explicit grounding to actual mental states. We propose self-supervised alignment strategies and present this work to solicit community feedback on grounding approaches.

心智理论分层模型隐变量导航推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。