在真实场景中用多模态数据自动评估人机互动亲密度,提升机器人自适应能力。
Multimodal Rapport Estimation in Real-World HRI

- 用62次日本药店真实交互数据,测试多模态模型对亲密度的预测能力。
- 零样本大模型表现优异,音频与视觉模型信息互补,融合模型效果最佳。
- 模型性能受互动时长和人数影响,需考虑真实场景的复杂性。
在真实人机交互(HRI)中评估互动质量是一项重要挑战。若能可靠估计互动质量,可优化对话策略,使机器人实现自主行为调整。然而,现有自动评估方法多基于受控实验室环境,尚不清楚是否可直接应用于真实场景——用户可能随时中断,多人参与也更自然出现。本研究基于62段在日本药房采集的多模态交互数据,探究第三方评分的亲密度自动估计问题。对比了零样本大语言模型(LLM)、预训练文本、音频和视觉模型及其预测级融合。结果表明,在真实场景中,零样本大模型表现强劲,音频与视觉模型提供互补信息。尤其Gemini 2.5 Flash作为单一模型表现突出,而结合Gemini(文本)、HuBERT(音频)与V-JEPA(视觉)的融合模型整体最优。进一步分析显示,模型表现随互动时长与参与人数变化而异。研究提示:真实场景中的亲密度评估需考虑实验室之外的上下文变异性,模型设计也应予以响应。
原文摘要 · Abstract (English)
Evaluating interaction quality in real-world HRI is an important challenge. If interaction quality can be estimated reliably, the results can be used to improve dialogue strategies and ultimately enable robots to adapt their behavior autonomously. However, existing automatic evaluation methods have been developed primarily in controlled laboratory settings, and it remains unclear whether they can be directly applied to real-world environments, where users are free to disengage and multi-party participation may arise naturally. In this study, we investigate the automatic estimation of third-party-rated rapport scores using 62 sessions of multimodal recordings collected in a Japanese drugstore. We compare zero-shot LLMs, pretrained text, audio, and visual models, and their prediction-level fusion. The results show that, in real-world HRI, zero-shot LLMs achieve strong performance, while audio and visual models tend to provide complementary information. In particular, Gemini 2.5 Flash performs strongly as a single model, and a fusion model combining Gemini (text) with HuBERT and V-JEPA performs best overall. Further analyses showed that estimation performance varied across interaction-duration and group-size conditions. These findings suggest that rapport estimation in real-world HRI requires evaluation and model design that account for contextual variability beyond that assumed in laboratory settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。