arXiv:2604.13581cs.CV2026-04

用视觉语言模型和几何约束,从单目视频重建多人交互动作

SocialMirror: Reconstructing 3D Human Interaction Behaviors from Monocular Videos with Semantic and Geometric Guidance

论文配图:SocialMirror: Reconstructing 3D Human Interaction Behaviors from Monocular Videos with Semantic and Geometric Guidance
图 1 · 摘自论文原文
  • 结合语义与几何信息,用扩散模型补全遮挡部位
  • 生成流畅无抖动的动作,保持身体接触和空间关系合理
  • 适合虚拟现实、人机协作等需要精准交互建模的场景

在增强现实中的真实虚拟互动、体育运动的精确动作分析以及人机协同任务中,准确重建近距离交互场景下的人体行为至关重要。然而,在单目视频中进行人体重建仍面临严重相互遮挡问题,导致局部动作模糊、时间连续性断裂和空间关系错误。本文提出 SocialMirror,一个基于扩散模型的框架,融合语义与几何引导信号以解决上述挑战。首先,利用视觉-语言模型生成的高层交互描述,指导语义引导的动作补全模块,恢复被遮挡的身体部分并消除局部姿态歧义。其次,提出一种序列级时序优化器,在采样过程中引入几何约束,确保动作平滑无抖动,并维持合理的身体接触与空间关系。在多个交互基准上的评估表明,SocialMirror 在重建交互人体网格方面达到当前最优性能,展现出对未见数据集及真实场景的强大泛化能力。代码将在发表后公开。

原文摘要 · Abstract (English)

Accurately reconstructing human behavior in close-interaction scenarios is crucial for enabling realistic virtual interactions in augmented reality, precise motion analysis in sports, and natural collaborative behavior in human-robot tasks. Reliable reconstruction in these contexts significantly enhances the realism and effectiveness of AI-driven interactive applications. However, human reconstruction from monocular videos in close-interaction scenarios remains challenging due to severe mutual occlusions, leading local motion ambiguity, disrupted temporal continuity and spatial relationship error. In this paper, we propose SocialMirror, a diffusion-based framework that integrates semantic and geometric cues to effectively address these issues. Specifically, we first leverage high-level interaction descriptions generated by a vision-language model to guide a semantic-guided motion infiller, hallucinating occluded bodies and resolving local pose ambiguities. Next, we propose a sequence-level temporal refiner that enforces smooth, jitter-free motions, while incorporating geometric constraints during sampling to ensure plausible contact and spatial relationships. Evaluations on multiple interaction benchmarks show that SocialMirror achieves state-of-the-art performance in reconstructing interactive human meshes, demonstrating strong generalization across unseen datasets and in-the-wild scenarios. The code will be released upon publication.

3D重建人体交互扩散模型单目视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。