arXiv:2506.16755cs.CLcs.AI2025-06EMNLP被引 10

用语言和视觉信息构建智能体模型,实现更贴近人类的社交推理。

Language-Informed Synthesis of Rational Agent Models for Grounded Theory-of-Mind Reasoning On-The-Fly

  • 融合语言与视觉输入,生成情境化智能体表征
  • 在多个认知科学任务中逼近人类判断,性能超越现有模型
  • 适合研究社交认知、人机交互及具身智能的学者

现实世界中的社会推断通常需要整合多模态信息。语言在社交场景中尤其关键,尤其是在陌生情境下,能提供环境动态的抽象描述和难以通过视觉观察到的个体具体信息。本文提出语言引导的理性智能体合成框架(LIRAS),通过融合语言与视觉输入进行上下文相关的社会推断。LIRAS将多模态社会推理建模为构建结构化但情境特定的智能体与环境表征的过程:利用多模态语言模型将语言与视觉输入解析为统一的符号表示,再通过贝叶斯逆规划引擎生成细粒度的概率判断。在源自认知科学实验的多个现有与新任务上,该模型(使用轻量级视觉语言模型实例化)在捕捉人类判断方面,优于消融实验与当前最优模型,在所有领域均表现更优。

原文摘要 · Abstract (English)

Drawing real world social inferences usually requires taking into account information from multiple modalities. Language is a particularly powerful source of information in social settings, especially in novel situations where language can provide both abstract information about the environment dynamics and concrete specifics about an agent that cannot be easily visually observed. In this paper, we propose Language-Informed Rational Agent Synthesis (LIRAS), a framework for drawing context-specific social inferences that integrate linguistic and visual inputs. LIRAS frames multimodal social reasoning as a process of constructing structured but situation-specific agent and environment representations - leveraging multimodal language models to parse language and visual inputs into unified symbolic representations, over which a Bayesian inverse planning engine can be run to produce granular probabilistic judgments. On a range of existing and new social reasoning tasks derived from cognitive science experiments, we find that our model (instantiated with a comparatively lightweight VLM) outperforms ablations and state-of-the-art models in capturing human judgments across all domains.

社会推理多模态贝叶斯推理语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。