用计算理性模型解释人类因记忆有限导致的非最优行为。
More Than Irrational: Modeling Belief-Biased Agents
- 构建认知受限下基于偏见信念的理性决策模型
- 仅需100步观测即可准确推断记忆容量等认知限制
- 适合开发能适应用户记忆短板的智能助手
尽管人工智能技术迅猛发展,预测和推断用户或人类合作者的非最优行为仍是关键挑战。许多此类行为并非源于非理性,而是认知能力受限与世界信念偏见下的理性选择。本文提出一类计算理性(CR)用户模型,用于建模在认知约束下、基于偏见信念进行最优决策的代理。核心创新在于显式建模有限记忆过程如何导致动态不一致且有偏差的信念状态,从而引发次优的序列决策。我们解决了从被动观察中实时识别用户特定认知约束并推断偏见信念状态的难题。对于具有显式参数化认知过程的CR模型族,该问题可解。为此,我们提出一种基于嵌套粒子滤波的高效在线推理方法,可同步追踪用户隐含信念状态并估计未知认知约束。在以记忆衰减为例的认知约束导航任务中验证:(1) 该模型能生成符合直觉的不同记忆容量下的行为;(2) 推理方法能从有限观测(≤100步)中准确高效恢复真实认知约束。进一步表明该方法为开发自适应AI助手机提供理论基础,实现考虑用户记忆局限的动态辅助。
原文摘要 · Abstract (English)
Despite the explosive growth of AI and the technologies built upon it, predicting and inferring the sub-optimal behavior of users or human collaborators remains a critical challenge. In many cases, such behaviors are not a result of irrationality, but rather a rational decision made given inherent cognitive bounds and biased beliefs about the world. In this paper, we formally introduce a class of computational-rational (CR) user models for cognitively-bounded agents acting optimally under biased beliefs. The key novelty lies in explicitly modeling how a bounded memory process leads to a dynamically inconsistent and biased belief state and, consequently, sub-optimal sequential decision-making. We address the challenge of identifying the latent user-specific bound and inferring biased belief states from passive observations on the fly. We argue that for our formalized CR model family with an explicit and parameterized cognitive process, this challenge is tractable. To support our claim, we propose an efficient online inference method based on nested particle filtering that simultaneously tracks the user's latent belief state and estimates the unknown cognitive bound from a stream of observed actions. We validate our approach in a representative navigation task using memory decay as an example of a cognitive bound. With simulations, we show that (1) our CR model generates intuitively plausible behaviors corresponding to different levels of memory capacity, and (2) our inference method accurately and efficiently recovers the ground-truth cognitive bounds from limited observations ($\le 100$ steps). We further demonstrate how this approach provides a principled foundation for developing adaptive AI assistants, enabling adaptive assistance that accounts for the user's memory limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。