构建首个两人动态交互大规模数据集,支持长时复杂互动建模。
InterAct: A Large-Scale Dataset of Dynamic, Expressive and Interactive Activities between Two People in Daily Scenarios
- 采集241段双人协作动作,含语音、姿态、表情多模态数据。
- 提出基于扩散模型的交互行为生成方法,提升面部与肢体动作准确性。
- 适合研究人机交互、具身智能及多角色行为生成的学者使用。
针对日常场景中两人交互行为的精准捕捉问题,现有工作多仅关注单人或对话手势,且假设参与者姿态位置基本不变。本文提出同时建模两人活动的新范式,聚焦目标驱动、动态且语义一致的长时交互。为此,我们构建了名为InterAct的多模态数据集,包含241个运动序列,每段持续1分钟以上,两名参与者扮演不同角色并标注情绪标签,协同完成任务或开展互动。数据涵盖双人语音、身体动作和面部表情。该数据集包含多样化复杂动作及此前罕见的长期交互模式。我们还提出一种简单有效的基于扩散模型的方法,可从语音输入生成两人交互的面部表情与身体动作。该方法分层回归身体动作,并引入新型微调机制提升唇部动作精度。相关数据与代码已公开于https://hku-cg.github.io/interact/。
原文摘要 · Abstract (English)
We address the problem of accurate capture of interactive behaviors between two people in daily scenarios. Most previous works either only consider one person or solely focus on conversational gestures of two people, assuming the body orientation and/or position of each actor are constant or barely change over each interaction. In contrast, we propose to simultaneously model two people's activities, and target objective-driven, dynamic, and semantically consistent interactions which often span longer duration and cover bigger space. To this end, we capture a new multi-modal dataset dubbed InterAct, which is composed of 241 motion sequences where two people perform a realistic and coherent scenario for one minute or longer over a complete interaction. For each sequence, two actors are assigned different roles and emotion labels, and collaborate to finish one task or conduct a common interaction activity. The audios, body motions, and facial expressions of both persons are captured. InterAct contains diverse and complex motions of individuals and interesting and relatively long-term interaction patterns barely seen before. We also demonstrate a simple yet effective diffusion-based method that estimates interactive face expressions and body motions of two people from speech inputs. Our method regresses the body motions in a hierarchical manner, and we also propose a novel fine-tuning mechanism to improve the lip accuracy of facial expressions. To facilitate further research, the data and code is made available at https://hku-cg.github.io/interact/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。