用外观和社交距离先验重建真实人体互动动作
Reconstructing Close Human Interaction with Appearance and Proxemics Reasoning
- 双分支优化框架结合外观与社交距离先验
- 在复杂场景下实现高精度互动动作估计
- 适合研究人体行为理解与视频生成的学者
由于视觉模糊和人物相互遮挡,现有姿态估计方法难以从真实场景视频中恢复合理的近距离互动。即使最先进的大模型(如SAM)也难以准确区分此类场景中的人体语义。本文发现人体外观可作为解决该问题的直接线索。基于此,提出一种双分支优化框架,通过人体外观、社交距离和物理规律约束,重建准确的交互动作。首先训练扩散模型学习人体社交行为与姿态先验;随后将训练好的网络与两个可优化张量融入双分支框架,联合优化人体动作与外观。设计了基于3D高斯、2D关键点和网格穿透的多重约束辅助优化。借助社交距离先验与多样化约束,方法可在复杂环境中准确估计真实互动。此外,构建了一个带有伪真值交互标注的数据集,有助于未来姿态估计与人体行为理解研究。多个基准测试结果表明,本方法优于现有方法。代码与数据已公开于https://www.buzhenhuang.com/works/CloseApp.html。
原文摘要 · Abstract (English)
Due to visual ambiguities and inter-person occlusions, existing human pose estimation methods cannot recover plausible close interactions from in-the-wild videos. Even state-of-the-art large foundation models~(\eg, SAM) cannot accurately distinguish human semantics in such challenging scenarios. In this work, we find that human appearance can provide a straightforward cue to address these obstacles. Based on this observation, we propose a dual-branch optimization framework to reconstruct accurate interactive motions with plausible body contacts constrained by human appearances, social proxemics, and physical laws. Specifically, we first train a diffusion model to learn the human proxemic behavior and pose prior knowledge. The trained network and two optimizable tensors are then incorporated into a dual-branch optimization framework to reconstruct human motions and appearances. Several constraints based on 3D Gaussians, 2D keypoints, and mesh penetrations are also designed to assist the optimization. With the proxemics prior and diverse constraints, our method is capable of estimating accurate interactions from in-the-wild videos captured in complex environments. We further build a dataset with pseudo ground-truth interaction annotations, which may promote future research on pose estimation and human behavior understanding. Experimental results on several benchmarks demonstrate that our method outperforms existing approaches. The code and data are available at https://www.buzhenhuang.com/works/CloseApp.html.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。