arXiv:2412.16698cs.CVcs.HC2024-12中稿 · ICME, 2025被引 3

实时预测用户互动意图、态度和动作,提升人机协作效率

Interact with me: Joint Egocentric Forecasting of Intent to Interact, Attitude and Social Actions

  • 基于1秒视频输入,用全身骨骼点构建图结构模型
  • 三项任务联合预测准确率达83.15%,支持实时推理
  • 适合需要主动交互的智能助手与机器人场景

为实现高效人机交互,代理应主动识别目标用户并预判交互内容。本文提出新任务:从代理的视角(第一人称)联合预测用户是否意图互动、对代理的态度及将执行的动作。为此,我们设计了SocialEgoNet——一种基于图的时空框架,通过分层多任务学习挖掘任务间依赖关系。该模型仅需1秒视频输入提取全身关键点(面部、手部、身体),实现高速推理。我们在现有第一人称人机交互数据集基础上,新增类别标签与边界框标注,构建新数据集JPL-Social。大量实验表明,模型可实现真正实时推理,各项任务平均准确率达83.15%,显著优于多个基线方法。额外标注数据与代码将在论文录用后公开。

原文摘要 · Abstract (English)

For efficient human-agent interaction, an agent should proactively recognize their target user and prepare for upcoming interactions. We formulate this challenging problem as the novel task of jointly forecasting a person's intent to interact with the agent, their attitude towards the agent and the action they will perform, from the agent's (egocentric) perspective. So we propose \emph{SocialEgoNet} - a graph-based spatiotemporal framework that exploits task dependencies through a hierarchical multitask learning approach. SocialEgoNet uses whole-body skeletons (keypoints from face, hands and body) extracted from only 1 second of video input for high inference speed. For evaluation, we augment an existing egocentric human-agent interaction dataset with new class labels and bounding box annotations. Extensive experiments on this augmented dataset, named JPL-Social, demonstrate \emph{real-time} inference and superior performance (average accuracy across all tasks: 83.15\%) of our model outperforming several competitive baselines. The additional annotations and code will be available upon acceptance.

人机交互第一人称视觉多任务学习实时预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。