端到端统一追踪多人人体网格,支持用户指令引导分析。
DETRAM: End-to-end DEtection, Tracking and Recovery of HumAn Meshes

- 用单一变换器解码器实现检测、重建与追踪一体化。
- 在多个数据集上达成顶尖追踪效果,支持用户指定人物追踪。
- 首次实现可提示的多人群体人体建模与追踪,适合交互式视频分析。
在人体网格恢复(HMR)任务中,多人群场景因人物众多且存在持续遮挡而极具挑战。尤其针对视频输入,需可靠一致地追踪每个个体。现有方法依赖预训练检测模块,增加运行时间并限制可追踪人数。本文提出DETRAM,一种统一框架,可同时完成多人群体人体的检测、重建与跨帧追踪,支持自动与用户提示两种方式。该框架采用单一变换器解码器,包含身份一致的可学习查询嵌入:检测查询发现新人物,追踪查询保持已有个体的姿态与形状,提示查询遵循用户指定身份。实验表明,DETRAM在PoseTrack21、3DPW、BEDLAM和MuPoTS-3D上均取得领先追踪性能,在BEDLAM与3DPW上重建精度也具竞争力。据我们所知,这是首个将可提示性、多人群体人体建模与追踪集成于端到端可训练框架中的方法,支持用户主导的视频人体分析。
原文摘要 · Abstract (English)
In the task of human mesh recovery (HMR), multi-person scenes are particularly difficult to handle due to the many entities that appear and occlusions between them over time. In particular for video inputs, there is a need to track each entity reliably and consistently. Existing methods rely on pretrained human detection modules, increasing their runtime and limiting the number of tracked entities. We present DETRAM, a unified framework for multi-person HMR and tracking that simultaneously detects, reconstructs, and tracks humans across time, both automatically and via user prompts. DETRAM uses a single transformer decoder with an identity-consistent set of learnable query embeddings that persist across frames: detection queries discover new people, tracking queries maintain pose and shape for existing individuals, and prompt queries follow user-specified identities. Our approach achieves state-of-the-art tracking results on PoseTrack21, 3DPW, BEDLAM, and MuPoTS-3D, and competitive reconstruction accuracy on BEDLAM and 3DPW, while uniquely supporting prompt-based tracking of individuals in multi-person scenes. To our knowledge, this is the first method to unify promptability and multi-person HMR with tracking in an end-to-end trainable framework, enabling user-directed human analysis in videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。