3D人体姿态特征能有效解释人类社会感知,优于主流视觉模型。
Simple 3D Pose Features Support Human and Machine Social Scene Understanding
- 用3D姿态点提取视频中人物空间位置与朝向信息
- 简化后的3D特征比2D特征更能预测人类社会判断
- 该方法可提升模型对社会场景的理解能力,适合具身智能研究
人类能从视觉输入中轻松识别社会互动,但其内在计算机制仍不明确,且最先进的深度神经网络(DNNs)在社交互动识别任务上仍面临挑战。我们假设人类依赖3D空间姿态信息进行社会判断,而这一信息在多数视觉DNN中缺失。为此,我们开发了一种新的姿态与深度估计流程,自动从短视频中提取3D身体关节位置。对比这些关节特征与超过350个视觉DNN的嵌入表示在预测人类社会判断上的表现,发现3D关节特征优于大多数DNN。进一步将3D关节压缩为仅包含人物3D位置与方向的极简特征集,发现该特征集虽简单,却是解释完整关节集性能的必要且充分条件。该最小3D特征集还能预测DNN与人类社会判断的一致性程度,并显著提升其在任务中的表现。结果表明,人类社会感知依赖于简单、显式的3D姿态信息。
原文摘要 · Abstract (English)
Humans effortlessly recognize social interactions from visual input, yet the underlying computations remain unknown, and social interaction recognition challenges even the most advanced deep neural networks (DNNs). Here, we hypothesized that humans rely on 3D visuospatial pose information to make social judgments, and that this information is largely absent from most vision DNNs. To test these hypotheses, we used a novel pose and depth estimation pipeline to automatically extract 3D body joint positions from short video clips. We compared the ability of these body joints to predict human social judgments in the videos with embeddings from over 350 vision DNNs. We found that body joints predicted social judgments better than most DNNs. We then reduced the 3D body joints to an even more compact feature set describing only the 3D position and direction of people in the videos. We found that this minimal 3D feature set, but not its 2D counterpart, was necessary and sufficient to explain the prediction performance of the full set of body joints. These minimal 3D features also predicted the extent to which DNNs aligned with human social judgments and significantly improved their performance on these tasks. Together, these findings demonstrate that human social perception depends on simple, explicit 3D pose information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。