arXiv:2601.19839cs.ROcs.AI2026-01

用大模型让机器人记住多人互动中的个性,实现长期个性化服务。

HARMONI: Multimodal Personalization of Multi-User Human-Robot Interactions with LLMs

  • 通过多模态感知与上下文建模,实时识别说话人并更新记忆。
  • 在养老院实测中,用户满意度和个性化准确率显著优于基线方法。
  • 适合需要长期人际交互的陪护机器人研发与部署。

现有机器人交互系统在多用户环境中缺乏持续个性化与动态适应机制,限制了实际应用效果。本文提出HARMONI,一种基于大语言模型的多模态个性化框架,支持社交助手机器人管理长期多用户交互。该框架包含四个模块:(i) 感知模块,识别活跃说话人并提取多模态输入;(ii) 世界建模模块,维护环境与短期对话上下文表示;(iii) 用户建模模块,持续更新特定用户的长期档案;(iv) 生成模块,生成符合上下文与伦理的回应。在四个数据集上进行广泛评估与消融实验,并在养老院场景开展真实用户研究,结果表明HARMONI能有效支持说话人识别、在线记忆更新与伦理对齐的个性化,在用户建模准确率、个性化质量与用户满意度方面均优于基线大模型方法。

原文摘要 · Abstract (English)

Existing human-robot interaction systems often lack mechanisms for sustained personalization and dynamic adaptation in multi-user environments, limiting their effectiveness in real-world deployments. We present HARMONI, a multimodal personalization framework that leverages large language models to enable socially assistive robots to manage long-term multi-user interactions. The framework integrates four key modules: (i) a perception module that identifies active speakers and extracts multimodal input; (ii) a world modeling module that maintains representations of the environment and short-term conversational context; (iii) a user modeling module that updates long-term speaker-specific profiles; and (iv) a generation module that produces contextually grounded and ethically informed responses. Through extensive evaluation and ablation studies on four datasets, as well as a real-world scenario-driven user-study in a nursing home environment, we demonstrate that HARMONI supports robust speaker identification, online memory updating, and ethically aligned personalization, outperforming baseline LLM-driven approaches in user modeling accuracy, personalization quality, and user satisfaction.

人机交互个性化大模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。