构建真实极端场景下的第一人称视角6D物体位姿数据集,推动鲁棒性模型研发。
EgoXtreme: A Dataset for Robust Object Pose Estimation in Egocentric Views under Extreme Conditions
- 从第一人称视角采集工业、体育、救援三类极端场景数据
- 在低光、强模糊、烟雾遮挡下,现有模型性能显著下降
- 证明仅靠图像修复无法提升极端条件表现,时间信息更关键
智能眼镜在双手忙碌、专注任务的场景中日益重要,第一人称视角下的6D物体位姿估计成为理解佩戴者上下文的关键。然而,现有基准难以模拟真实应用中的严重运动模糊、动态光照和视觉遮挡,导致实验室数据与现实场景差距显著。为此,我们提出EgoXtreme,一个大规模第一人称视角6D位姿估计数据集,涵盖工业维护、体育运动和应急救援三类挑战性场景,通过极端光照、剧烈运动模糊和烟雾引入严重感知歧义。对主流可泛化位姿估计算法在该数据集上的评估显示,其泛化能力在极端条件下失效,尤其在低光环境下表现更差。进一步实验表明,单纯应用图像恢复(如去模糊)无法带来正向改进;而基于追踪的方法在高速运动场景中表现更好,说明利用时间信息具有重要意义。结论是,EgoXtreme是开发和评估下一代真实世界第一人称视觉鲁棒位姿模型的必备资源。数据集与代码已公开于https://taegyoun88.github.io/EgoXtreme/
原文摘要 · Abstract (English)
Smart glass is emerging as an useful device since it provides plenty of insights under hands-busy, eyes-on-task situations. To understand the context of the wearer, 6D object pose estimation in egocentric view is becoming essential. However, existing 6D object pose estimation benchmarks fail to capture the challenges of real-world egocentric applications, which are often dominated by severe motion blur, dynamic illumination, and visual obstructions. This discrepancy creates a significant gap between controlled lab data and chaotic real-world application. To bridge this gap, we introduce EgoXtreme, a new large-scale 6D pose estimation dataset captured entirely from an egocentric perspective. EgoXtreme features three challenging scenarios - industrial maintenance, sports, and emergency rescue - designed to introduce severe perceptual ambiguities through extreme lighting, heavy motion blur, and smoke. Evaluations of state-of-the-art generalizable pose estimators on EgoXtreme indicate that their generalization fails to hold in extreme conditions, especially under low light. We further demonstrate that simply applying image restoration (e.g., deblurring) offers no positive improvement for extreme conditions. While performance gain has appeared in tracking-based approach, implying using temporal information in fast-motion scenarios is meaningful. We conclude that EgoXtreme is an essential resource for developing and evaluating the next generation of pose estimation models robust enough for real-world egocentric vision. The dataset and code are available at https://taegyoun88.github.io/EgoXtreme/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。