提升头戴设备姿态估计的置信区间可靠性,解决困难帧覆盖不足问题。
Adaptive Geodesic Conformal Prediction for Egocentric Camera Pose Estimation
- 用测地线SE(3)度量难度,更准确识别物理上难预测的帧
- 在最难25%帧上覆盖率从60%提升至93%,整体保持90%目标
- 无需测试时图像,跨用户迁移,适合AR和辅助设备部署
头戴式设备的增强现实与辅助应用不仅需要高精度姿态估计,还需保证不确定性区域。基于校准预测(CP)可在不重新训练的情况下提供此类保障,但我们发现标准CP使用单一固定阈值时,整体覆盖率达90%,但在最困难的25%帧(Q4)中覆盖率仅约60%,存在约30个百分点的条件覆盖差距,该现象在12名参与者、3种预测器和3个时间窗口(共108次评估)下均一致出现在EPIC-Fields数据集上。我们进一步表明,采用测地线SE(3)非一致性评分可比欧氏距离更有效地识别物理上更难的帧,其Q4帧重叠率仅为15%-26%,且测地线Q4帧的真实相机位移高出2-3倍。为弥合覆盖差距,我们提出DINOv2-Bridge自适应CP:一种两阶段难度估计算法,仅需单个源参与者的训练数据,即可在测试时无图像情况下实现跨用户迁移,将Q4覆盖率从约0.75提升至约0.93,同时维持整体覆盖率在90%目标水平。
原文摘要 · Abstract (English)
Egocentric pose estimation for Augmented Reality (AR) and assistive devices requires not just accurate predictions but guaranteed uncertainty regions. Conformal prediction (CP) provides such guarantees without retraining, but we show that standard CP with a single fixed threshold achieves nominal 90% overall coverage while covering only ~60% of the hardest 25% of frames (Q4) -- a ~30 percentage-point conditional coverage gap consistent across 12 participants, 3 predictors, and 3 horizons (108 evaluations) on EPIC-Fields. We further show that a geodesic SE(3) nonconformity score identifies physically harder frames than Euclidean scoring, with only 15-26% Q4 overlap and 2-3x higher ground-truth camera displacement for geodesic Q4 frames. To close the coverage gap, we propose DINOv2-Bridge adaptive CP: a two-stage difficulty estimator trained on a single source participant that transfers cross-participant without any images at test time, improving Q4 coverage from ~0.75 to ~0.93 while maintaining overall coverage at the 90% target.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。