用时间运动信息增强立体视觉,提升第一视角3D人体姿态估计鲁棒性
TSR-Ego: Temporally Guided Stereo Refinement Framework for Egocentric 3D Human Pose Estimation

- 通过时序卷积融合历史视觉证据,指导特征采样
- 在真实场景下相对基线提升12.3%的3D姿态精度
- 适合需要稳定在线推理的第一视角应用
从头戴式双目相机获取的第一视角3D人体姿态估计面临鱼眼畸变、严重自遮挡及关节点频繁超出视场的问题。现有方法虽通过热图提升、立体匹配和基于Transformer的优化改进性能,但多依赖当前帧局部线索,仅将时序信息作为辅助上下文,导致在当前帧立体线索弱、被遮挡或模糊时表现受限。本文提出TSR-Ego,一种时序引导的立体精修框架,将短期运动证据与投影引导的特征采样相结合。模型首先通过因果深度可分离时序卷积丰富密集立体特征图,使历史视觉信息影响特征空间后再进行可变形交叉注意力;单阶段因果立体解码器则通过时序自注意力、关节自注意力和鱼眼可变形立体交叉注意力,利用动态姿态估计生成2D采样参考,精修3D关节查询。与多数在姿态预测后才引入时序推理的方法不同,TSR-Ego在特征采样和关节表示构建阶段即利用运动上下文,同时保持无需未来帧的在线推理能力。在UnrealEgo2和UnrealEgo-RW数据集上的实验表明,该方法达到最新性能,尤其在真实世界序列上提升显著。
原文摘要 · Abstract (English)
Egocentric 3D human pose estimation from head-mounted stereo cameras is challenging due to fisheye distortion, severe self-occlusion, and frequent truncation of body joints outside the camera field of view. Recent stereo egocentric methods have improved performance through heatmap lifting, stereo correspondence, and transformer-based refinement, but they often rely heavily on frame-local evidence or use temporal information only as auxiliary pose-level context. This limits robustness when current-frame stereo cues are weak, occluded, or ambiguous. We propose TSR-Ego, a temporally guided stereo framework that couples short-term motion evidence with projection-guided feature sampling. The model first enriches dense stereo feature maps using a causal depthwise-separable temporal convolution, allowing past visual evidence to influence the feature space before deformable cross-attention. A single-stage causal stereo decoder then refines learned 3D joint queries through temporal self-attention, joint self-attention, and fisheye deformable stereo cross-attention, using the evolving pose estimate to generate 2D sampling references. Unlike methods that apply temporal reasoning mainly after pose prediction, TSR-Ego uses motion context to shape both the sampled stereo features and the joint representations while preserving online inference without future frames. Experiments on UnrealEgo2 and UnrealEgo-RW show state-of-the-art performance, with especially strong gains on real-world sequences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。