arXiv:2606.04480cs.CVcs.HC2026-06

IMPose通过动态传播修正,高效实现多人动态姿态标注。

IMPose: Interactive Multi-person Pose Estimation with Dynamic Correction Propagation

论文配图:IMPose: Interactive Multi-person Pose Estimation with Dynamic Correction Propagation
图 1 · 摘自论文原文
  • 双层追踪机制:关键点级与实例级协同保证跨帧一致性。
  • 每1050帧仅需27次点击,84帧视频每轨迹3次点击,效率显著提升。
  • 适合需要低交互标注的多人动作分析场景,尤其在遮挡和模糊下稳定。

高质量动态人体姿态标注为人工智能提供精确运动学信息以掌握人类行为,但依然费时费力。现有标注工具或缺乏时间修正传播,或在多人场景中表现不佳,需大量人工干预。本文提出IMPose,一种交互式多人动态姿态标注工具。其采用双层追踪机制,将标注者的一帧多人姿态修正跨视频传播。关键点级通过序列建模实现修正的时间传播,实例级则利用关键点感知嵌入与相对位置编码维持多人群体跨帧一致性。为进一步提升鲁棒性,IMPose在轨迹库中保存历史姿态与实例线索,增强长程时间关联,稳定遮挡和运动模糊等复杂情况下的标注。该框架将稀疏的人工修正转化为密集连贯的姿态轨迹,大幅减少帧间重复修正。大量实验表明,IMPose在不同交互预算下均保持良好精度-效率平衡,尤其在低点击设置中优势明显:在3DPW上每1050帧仅需27次点击,在PoseTrack21上每轨迹每84帧仅需3次点击。我们还以10名标注员10小时的代价扩展了PoseTrack21,新增18.8万姿态实例(355万关键点)。工具代码与扩展数据集将开源。

原文摘要 · Abstract (English)

High-quality dynamic human pose annotation equips AI with precise motion kinematics to enable human behavior mastery, yet remains labor-intensive and time-consuming. Current annotation tools either lack temporal correction propagation or fail in multi-person scenarios, necessitating excessive manual intervention. In this paper, we introduce IMPose, an interactive tool for multi-person dynamic pose annotation. It features a dual-level tracking mechanism that propagates one-frame multi-person pose corrections from annotators across entire videos. The keypoint-level ensures corrections temporal propagation via sequential modeling, while the instance-level employs keypoint-aware embedding with relative positional encoding to maintain multi-person cross-frame consistency. To further improve robustness, IMPose maintains historical pose and instance cues in a trajectory bank, which enhances long-range temporal association and stabilizes annotation in challenging cases such as occlusion and motion blur. By converting sparse human corrections into dense and coherent pose trajectories, our framework significantly reduces repeated manual refinement across frames. Extensive experiments show that IMPose consistently achieves a strong accuracy efficiency trade off under different interaction budgets, demonstrating particular advantages in low click annotation settings. IMPose achieves high precision annotation with high efficiency, requiring only 27 clicks per 1,050 frame video on 3DPW and 3 clicks per tracklet per 84-frame on PoseTrack21. We further expand PoseTrack21 with 188K pose instances (3.55M keypoints) at a minimal cost of 10 annotators in 10 hours. The annotation tool, codes, and extended dataset will be open-sourced.

姿态估计交互标注多人跟踪高效标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。