arXiv:2506.14212cs.AI2025-06被引 1

用多模态线索推理看不见的物体,模型表现接近人类。

What's in the Box? Reasoning about Unseen Objects from Multimodal Cues

  • 结合神经网络与贝叶斯推理,融合视觉、听觉和语言信息
  • 在猜盒子游戏测试中,模型与人类判断相关性达0.82
  • 适合研究认知推理、人机共情或跨模态理解的学者

人们常通过整合听觉、视觉、语言及先验知识等多源信息,推断无法直接观察的物体。本文提出一种神经符号模型:利用神经网络解析开放式的多模态输入,并通过贝叶斯模型融合不同信息源以评估多种假设。我们在新设计的“盒子里有什么?”物体猜测游戏中评估该模型,参与者观看实验者摇动盒子的视频后猜测内部物体。通过人类实验发现,该模型与人类判断的相关性高达0.82,而单模态消融模型和大型多模态神经网络基线模型的相关性较差。

原文摘要 · Abstract (English)

People regularly make inferences about objects in the world that they cannot see by flexibly integrating information from multiple sources: auditory and visual cues, language, and our prior beliefs and knowledge about the scene. How are we able to so flexibly integrate many sources of information to make sense of the world around us, even if we have no direct knowledge? In this work, we propose a neurosymbolic model that uses neural networks to parse open-ended multimodal inputs and then applies a Bayesian model to integrate different sources of information to evaluate different hypotheses. We evaluate our model with a novel object guessing game called ``What's in the Box?'' where humans and models watch a video clip of an experimenter shaking boxes and then try to guess the objects inside the boxes. Through a human experiment, we show that our model correlates strongly with human judgments, whereas unimodal ablated models and large multimodal neural model baselines show poor correlation.

多模态推理贝叶斯建模认知模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。