arXiv:2505.06278cs.ROcs.HC2025-05中稿 · ACM Multimedia 202…被引 2

用人体姿态实现高效鲁棒的人机社交理解,抗干扰能力强。

Robust Understanding of Human-Robot Social Interactions through Multimodal Distillation

  • 通过多模态知识蒸馏,让小模型仅凭身体姿态理解社交行为。
  • 在输入51%被破坏时仍比基线高14.75%准确率,且模型小于1%参数量。
  • 适合部署在资源受限的实时人机交互场景,如服务机器人。

随着社交机器人与智能体需求增长,其需从自身视角分析社交场景与行为线索。现有研究稀少且计算开销大,难以实时部署或在信息有限时表现不佳。本文提出一种多模态知识蒸馏框架:教师模型融合身体、面部、手势、注视及图像等多模态输入,将知识迁移至仅依赖身体姿态的学生模型。在两个公开的人机交互数据集上验证,学生模型在多个下游社交理解任务中平均准确率提升14.75%,即使输入高达51%被破坏仍保持性能;模型参数量不足教师模型的1%,延迟仅为教师的11.9%。代码与数据已开源。

原文摘要 · Abstract (English)

There is a growing need for social robots and intelligent agents that can effectively interact with and support users. For the interactions to be seamless, the agents need to analyse social scenes and behavioural cues from their (robot's) perspective. Works that model human-agent interactions in social situations are few; and even those existing ones are computationally too intensive to be deployed in real time or perform poorly in real-world scenarios when only limited information is available. We propose a knowledge distillation framework that models social interactions through various multimodal cues, and yet is robust against incomplete and noisy information during inference. We train a teacher model with multimodal input (body, face and hand gestures, gaze, raw images) that transfers knowledge to a student model which relies solely on body pose. Extensive experiments on two publicly available human-robot interaction datasets demonstrate that our student model achieves an average accuracy gain of 14.75% over competitive baselines on multiple downstream social understanding tasks, even with up to 51% of its input being corrupted. The student model is also highly efficient - less than 1% in size of the teacher model in terms of parameters and its latency is 11.9% of the teacher model. Our code and related data are available at github.com/biantongfei/SocialEgoMobile.

人机交互知识蒸馏多模态机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。