arXiv:2609.07394cs.CV2026-09

人类比机器更懂人机互动意图,尤其在视频中。

Social Intuition vs. Machine Reasoning: Anticipating Human-Robot Interaction from multiple modalities

论文配图:Social Intuition vs. Machine Reasoning: Anticipating Human-Robot Interaction from multiple modalities
图 1 · 摘自论文原文
  • 用姿态或视频预测人类是否互动,人类表现优于轻量模型
  • 完整视角视频下人类F1得分比视觉语言模型高0.2
  • 模型大小不影响性能,说明单纯推理不够,需社会直觉

预测从服务机器人视角出发,一个人是否会与之互动,是人类凭借多种线索的直观判断。我们研究了人类在仅用姿态或完整视频输入时的表现,并对比了轻量级姿态模型和最先进的视觉-语言模型(VLMs)。实验基于HUI360数据集的100条测试轨迹(25条正例,75条负例)进行。结果表明:仅使用姿态输入时,人类虽优于轻量模型,但差距不大(F1-score提升+0.08);而提供包含目标边框的全景视频时,人类显著优于所有VLMs(F1-score提升+0.2)。此外,不同尺寸的VLMs在不同输入条件下的表现并不随规模增长而提升。研究证实,互动预测对社交机器人仍具挑战性,仅具备推理能力的模型难以超越人类的社会直觉。

原文摘要 · Abstract (English)

Anticipating whether a person will interact from one's own perspective is a highly intuitive task for humans, that relies on a combination of cues. We investigate how humans perform at predicting a person's intention to interact from a service robot's point of view, using pose-only or full video input, then benchmark different lightweight pose-based models and state-of-the-art vision-language models. We conducted our benchmark on the HUI360 dataset on a fixed pilot subset of 100 test tracks (25 positive, 75 negative). We found that with pose-only input, human annotators outperform lightweight trained pose models but not by large margins (+0.08 in F1-Score). But when given full egocentric video with a target bounding box, human annotators perform substantially better and largely outperform the Vision-Language Models (+0.2 in F1-Score). We also compared VLMs of different size and under different input conditions, and found that the best results do not correlate with model size. Our result confirms that predicting interactions is a challenging task for social robots and that reasoning-capable models are necessary but their actual reasoning capabilities alone do not suffice to match the social intuition of humans.

人机交互社会直觉视觉语言模型姿态估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。