arXiv:2606.05896cs.CV2026-06

让虚拟人具备共情能力,能推理他人想法并实时回应。

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind

论文配图:Resonant Minds: Closed-Loop Social Avatars with Theory of Mind
图 1 · 摘自论文原文
  • 双代理闭环框架:感知-推理-表达连续互动
  • 在信息不对称下对话质量超越全知模式
  • 支持心理驱动的虚拟人行为,适合社交模拟研究

构建具备真实社交智能的数字人类,需统一认知推理与多模态生成。现有方法将二者割裂:大语言模型擅长对话但缺乏具身表达,扩散模型虽视觉逼真却忽略社会认知。为此,我们提出闭环双代理框架,整合感知、社会推理与表达于持续交互循环中。感知模块从视频分析对方多模态行为,社会推理模块通过心智理论推断隐藏心理状态,并以集成机制选择回应;表达模块生成可控制情绪的视频,同步合成说话人语音与面部表情及听者反应行为,捕捉双向动态。我们还构建了分层人格-场景数据集,包含心理学基础的人格设定与私密社交目标,用于信息不对称下的评估。实验表明,该方法在对话质量与视频生成指标上表现优异,尤其在关键对话维度超越全知脚本模式,表明在不确定性下显式推断心理状态可激发更深刻的对话。项目页:https://resonantminds.github.io/

原文摘要 · Abstract (English)

Creating lifelike digital humans with genuine social intelligence requires unifying cognitive reasoning and multimodal generation within a coherent framework. Current approaches treat these as separate tasks: Large Language Models excel at dialogue but lack embodied expression, while diffusion-based talking head models achieve visual fidelity but ignore social cognition. To bridge this gap, we propose a closed-loop dual-agent framework integrating perception, social reasoning, and expression into a continuous interaction cycle. The perception module analyzes partners' multimodal behaviors from video, while the social reasoning module infers hidden mental states through Theory of Mind and selects responses via an ensemble mechanism. The expression module then generates emotion-controllable videos that jointly synthesize speaker speech and facial expressions with listener reactive behaviors, capturing bidirectional dynamics absent in prior work. We further construct a hierarchical Persona-Scenario dataset with psychologically grounded personas and private social goals to support evaluation under information asymmetry. Experiments on this dataset demonstrate competitive or superior performance on both dialogue quality and video generation metrics. Notably, our method surpasses even the full-information Script mode on key dialogue quality dimensions, suggesting that explicit mental state inference under uncertainty can elicit more thoughtful dialogue than unrestricted information access. Project page: https://resonantminds.github.io/.

虚拟人社会智能心智理论多模态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。