用贝叶斯逐步更新实现多模态心智推理,让小模型指导大模型更准确理解他人心理。
Overcoming Multi-step Complexity in Multimodal Theory-of-Mind Reasoning: A Scalable Bayesian Planner
- 将心智推理拆解为逐步贝叶斯更新,避免复杂流程依赖
- 在多模态基准上比现有方法提升4.6%准确率,包括未见场景
- 适合需要理解复杂社会行为的AI系统研发者
心智理论(ToM)使人能够推断信念、欲望和意图等心理状态,是社会认知的基础。然而,现有计算式ToM方法依赖特定先验或深度微调,难以在多模态环境中扩展,且任务复杂度上升时泛化能力下降。为此,我们提出一种可扩展的贝叶斯ToM规划器,将ToM推理分解为分步贝叶斯更新。该框架引入弱到强控制机制,使小型语言模型(LM)专注于特定的似然估计,并将其推理行为迁移至大型模型(7B至405B)中,融合社会与世界知识。此协同策略使大模型对人类心理状态的推理符合贝叶斯原则。大量实验表明,该方法在多模态ToM基准上相较最先进方法提升4.6%准确率,包括具有挑战性的未见场景,从而建立复杂环境中的心理建模新标准。
原文摘要 · Abstract (English)
Theory-of-Mind (ToM) enables humans to infer mental states-such as beliefs, desires, and intentions-forming the foundation of social cognition. However, existing computational ToM methods rely on structured workflows with ToM-specific priors or deep model fine-tuning, which struggle with scalability in multimodal environments and fail to generalize as task complexity increases. To address these limitations, we propose a scalable Bayesian ToM planner that decomposes ToM reasoning into stepwise Bayesian updates. Our framework introduces weak-to-strong control, allowing smaller language models (LMs) to specialize in ToM-specific likelihood estimation and transfer their reasoning behaviors to larger LMs (7B to 405B) for integration with social and world knowledge. This synergistic approach aligns large-model inference of human mental states with Bayesian principles. Extensive experiments show that our method achieves a 4.6% accuracy improvement over state-of-the-art techniques on multimodal ToM benchmarks, including challenging unseen scenarios, thereby establishing a new standard for modeling human mental states in complex environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。