arXiv:2507.05751cs.CV2025-07被引 2

首个跨环境与传感器变化的6D姿态估计基准,揭示现有模型鲁棒性短板。

SenseShift6D: Multimodal RGB-D Benchmarking for Robust 6D Pose Estimation across Environment and Sensor Variations

  • 构建包含13种曝光、9种增益等1380种组合的RGB-D数据集
  • 实测主流模型在光照/传感器变化下性能下降超16.7个百分点
  • 适合关注真实场景部署与多模态自适应的研究者

近期6D物体姿态估计在LM-O、YCB-V和T-Less等基准上取得优异表现,但这些数据集均在固定光照与相机设置下采集,未涵盖真实世界中光照、曝光、增益或深度传感器模式的变化影响。为此,我们提出SenseShift6D,首个通过物理实验覆盖13种RGB曝光、9种RGB增益、自动曝光、4种深度捕获模式及5种光照水平的RGB-D数据集。针对6种常见家居物体,共获取198.8k RGB图像与20.0k深度图像(总计795.4k RGB-D场景),每种物体姿态对应1,380种独特的传感器-光照组合。对先进预训练通用姿态估计算法的实验显示,其在不同光照与传感器条件下性能显著波动,即便在相同物体与背景下训练测试的实例级模型也对环境与传感器变化高度敏感。这些发现表明,传感器与环境感知的鲁棒性是实际部署中亟待重视却尚未充分探索的方向。为展示该基准的价值,我们评估了无需重训练的测试时多模态传感器选择策略:理想化(即最优)控制器可带来最高+16.7个百分点的性能提升,而实用的基于一致性的代理方法仅实现轻微改善,凸显巨大优化空间与未来研究必要性。

原文摘要 · Abstract (English)

Recent advances on 6D object pose estimation have achieved high performance on representative benchmarks such as LM-O, YCB-V, and T-Less. However, these datasets were captured under fixed illumination and camera settings, leaving the impact of real-world variations in illumination, exposure, gain or depth-sensor mode largely unexplored. To bridge this gap, we introduce SenseShift6D, the first RGB-D dataset that physically sweeps 13 RGB exposures, 9 RGB gains, auto-exposure, 4 depth-capture modes, and 5 illumination levels. For six common household objects, we acquire 198.8k RGB and 20.0k depth images (i.e., 795.4k RGB-D scenes), providing 1,380 unique sensor-lighting permutations per object pose. Experiments with state-of-the-art pretrained, generalizable pose estimators reveal substantial performance variation across lighting and sensor settings, despite their large-scale pretraining. Strikingly, even instance-level estimators-trained and tested on identical objects and backgrounds-exhibit pronounced sensitivity to environmental and sensor shifts. These findings establish sensor- and environment-aware robustness as an underexplored yet essential dimension for real-world deployment, and motivate SenseShift6D as a necessary benchmark for the community. Finally, to illustrate the opportunity enabled by this benchmark, we evaluate test-time multimodal sensor selection without retraining. An idealized (oracle) controller yields remarkable gains of up to +16.7 pp for generalizable models, whereas a practical consistency-based proxy improves performance only marginally, highlighting substantial headroom and the need for future research on reliable sensor-aware adaptation.

6D姿态估计多模态感知鲁棒性评估数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。