arXiv:2604.12081cs.AI2026-04

让机器人像人一样有选择地记住重要社交经历,提升对话自然度。

Human-Inspired Context-Selective Multimodal Memory for Social Robots

论文配图:Human-Inspired Context-Selective Multimodal Memory for Social Robots
图 1 · 摘自论文原文
  • 模仿人类记忆机制,只存储高情绪或新奇场景的图文记忆。
  • 在社交场景数据上相关性达0.506,优于人类一致性(0.415)。
  • 适合需要长期个性化互动的社交机器人研发者。

记忆是社会互动的基础,使人能基于上下文回忆过往经历并调整行为。然而,当前多数社交机器人和具身智能体依赖非选择性的文本记忆,限制了个性化与情境感知能力。受认知神经科学启发,我们提出一种情境选择性、多模态记忆架构,可捕捉并检索文本与视觉情景记忆,优先保留高情绪显著性或场景新颖性的时刻。通过将记忆关联至特定用户,系统实现社会个性化回溯,支持更自然、有根基的对话。在自建社交场景数据集上评估选择性存储机制,获得Spearman相关系数0.506,超越人类一致性(ρ=0.415),且优于现有图像可记性模型。多模态检索实验中,融合方法使Recall@1最高提升13%。运行时评估表明系统保持实时性能。定性分析显示,该框架生成的回复更丰富、更具社会相关性。本工作通过融合人类启发的选择性与多模态检索,推动了社交机器人长时个性化交互的记忆设计。

原文摘要 · Abstract (English)

Memory is fundamental to social interaction, enabling humans to recall meaningful past experiences and adapt their behavior accordingly based on the context. However, most current social robots and embodied agents rely on non-selective, text-based memory, limiting their ability to support personalized, context-aware interactions. Drawing inspiration from cognitive neuroscience, we propose a context-selective, multimodal memory architecture for social robots that captures and retrieves both textual and visual episodic traces, prioritizing moments characterized by high emotional salience or scene novelty. By associating these memories with individual users, our system enables socially personalized recall and more natural, grounded dialogue. We evaluate the selective storage mechanism using a curated dataset of social scenarios, achieving a Spearman correlation of 0.506, surpassing human consistency ($ρ=0.415$) and outperforming existing image memorability models. In multimodal retrieval experiments, our fusion approach improves Recall@1 by up to 13\% over unimodal text or image retrieval. Runtime evaluations confirm that the system maintains real-time performance. Qualitative analyses further demonstrate that the proposed framework produces richer and more socially relevant responses than baseline models. This work advances memory design for social robots by bridging human-inspired selectivity and multimodal retrieval to enhance long-term, personalized human-robot interaction.

社交机器人多模态记忆情境选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。