arXiv:2507.23544cs.ROcs.CV2025-07中稿 · presentation at IE…

用多模态信号和多实例学习,更全面地评估人机交互中的用户体验。

User Experience Estimation in Human-Robot Interaction Via Multi-Instance Learning of Multimodal Social Signals

  • 融合面部表情与语音信号,通过多实例学习捕捉长期交互模式。
  • 在自建数据集上表现超越第三方人工评估者。
  • 适合关注人机交互体验评估的机器人研究者。

近年来,社交机器人需求增长,需根据用户状态调整行为。准确评估人机交互(HRI)中的用户体验(UX)是实现这一目标的关键。UX 是包含情感、参与度等多维度的指标,现有方法常仅关注单一维度。本研究提出一种基于多模态社会信号的 HRI UX 估计方法。我们构建了一个 UX 数据集,并开发了基于 Transformer 的模型,利用面部表情与语音进行估计。不同于依赖瞬时观测的传统模型,本方法采用多实例学习框架,捕捉短期与长期交互模式,从而更好地表征用户体验的动态变化。实验结果表明,该方法在 UX 估计上优于第三方人工评估者。

原文摘要 · Abstract (English)

In recent years, the demand for social robots has grown, requiring them to adapt their behaviors based on users' states. Accurately assessing user experience (UX) in human-robot interaction (HRI) is crucial for achieving this adaptability. UX is a multi-faceted measure encompassing aspects such as sentiment and engagement, yet existing methods often focus on these individually. This study proposes a UX estimation method for HRI by leveraging multimodal social signals. We construct a UX dataset and develop a Transformer-based model that utilizes facial expressions and voice for estimation. Unlike conventional models that rely on momentary observations, our approach captures both short- and long-term interaction patterns using a multi-instance learning framework. This enables the model to capture temporal dynamics in UX, providing a more holistic representation. Experimental results demonstrate that our method outperforms third-party human evaluators in UX estimation.

人机交互用户体验多模态深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。