arXiv:2601.03389cs.AIcs.LG2026-01中稿 · AAAI

让智能体学会自我觉察,用‘疼痛信念’提升学习能力

Exploration Through Introspection: A Self-Aware Reward Model

  • 用隐马尔可夫模型推断智能体的‘疼痛信念’作为内在信号
  • 自省型智能体在网格环境中表现远超基线模型
  • 可模拟人类复杂行为,适合研究意识与决策机制

理解人工智能体如何建模内部心理状态,是推进AI理论心智发展的关键。现有证据表明,自我意识与他者意识可能共享统一机制。本文通过强化学习智能体在网格世界中推断自身内部状态,提出一种受生物疼痛启发的自省探索机制,利用隐马尔可夫模型从在线观测中推断‘疼痛信念’,并将其融入主观奖励函数,研究自省对学习能力的影响。进一步,该框架用于对比正常与慢性疼痛感知模型的表现差异。结果表明,具备自省能力的智能体整体显著优于标准基线模型,且能复现复杂的人类行为特征。

原文摘要 · Abstract (English)

Understanding how artificial agents model internal mental states is central to advancing Theory of Mind in AI. Evidence points to a unified system for self- and other-awareness. We explore this self-awareness by having reinforcement learning agents infer their own internal states in gridworld environments. Specifically, we introduce an introspective exploration component that is inspired by biological pain as a learning signal by utilizing a hidden Markov model to infer "pain-belief" from online observations. This signal is integrated into a subjective reward function to study how self-awareness affects the agent's learning abilities. Further, we use this computational framework to investigate the difference in performance between normal and chronic pain perception models. Results show that introspective agents in general significantly outperform standard baseline agents and can replicate complex human-like behaviors.

自省学习强化学习理论心智隐马尔可夫

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。