arXiv:2603.14435cs.CV2026-03被引 1

用时空变压器实现单目4D人物交互实时重建

End-to-End Spatial-Temporal Transformer for Real-time 4D HOI Reconstruction

  • 端到端时空变换器,直接从视频和3D模板预测运动
  • 31.5帧/秒,比之前方法快600倍以上且精度更高
  • 适合需要实时高精度动作重建的应用场景

单目4D人体-物体交互(HOI)重建——从单个RGB视频中恢复动态人体与操作物体——因深度模糊和频繁遮挡仍具挑战。现有方法多依赖多阶段流程或迭代优化,导致推理延迟高,难以满足实时需求,且易累积误差。为此,我们提出THO:一种端到端的时空变换器,能从前向输入视频和3D模板中直接预测人体与协同物体的运动。THO通过利用时空HOI元组先验实现:空间先验利用接触区域邻近性,从人体线索推断被遮挡物体特征;时间先验捕捉跨帧运动相关性,优化物体表征并保证物理一致性。大量实验表明,THO在单块RTX 4090 GPU上达到31.5 FPS的推理速度,相比先前基于优化的方法提速超过600倍,同时提升重建精度与时间一致性。

原文摘要 · Abstract (English)

Monocular 4D human-object interaction (HOI) reconstruction - recovering a moving human and a manipulated object from a single RGB video - remains challenging due to depth ambiguity and frequent occlusions. Existing methods often rely on multi-stage pipelines or iterative optimization, leading to high inference latency, failing to meet real-time requirements, and susceptibility to error accumulation. To address these limitations, we propose THO, an end-to-end Spatial-Temporal Transformer that predicts human motion and coordinated object motion in a forward fashion from the given video and 3D template. THO achieves this by leveraging spatial-temporal HOI tuple priors. Spatial priors exploit contact-region proximity to infer occluded object features from human cues, while temporal priors capture cross-frame kinematic correlations to refine object representations and enforce physical coherence. Extensive experiments demonstrate that THO operates at an inference speed of 31.5 FPS on a single RTX 4090 GPU, achieving a >600x speedup over prior optimization-based methods while simultaneously improving reconstruction accuracy and temporal consistency. The project page is available at: https://nianheng.github.io/THO-project/

4D重建时空建模实时推理人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。