用15个像素点实现物体6维位姿精准估计,突破低成本传感器限制
Recovering Parametric Scenes from Very Few Time-of-Flight Pixels
- 结合前馈预测与可微渲染,从稀疏测距数据中重建场景参数
- 仅需15个像素即可在模拟和真实场景中恢复已知物体的6维位姿
- 适用于低分辨率、宽视场的消费级飞行时间传感器,适合嵌入式应用
本文旨在利用低成本商用飞行时间(ToF)传感器的极低空间分辨率(单像素)但高时间分辨光子计数数据,恢复三维参数化场景的几何结构。这类传感器每像素覆盖大视场,且时间分辨数据蕴含丰富场景信息,使简单场景的稀疏测量重建成为可能。我们探索了利用分布式少量测量(如仅15个像素)恢复具有强先验的简单参数化场景(如已知物体的6维位姿)的可行性。为此,设计了一种方法:结合前馈预测推断场景参数,并在分析-合成框架内使用可微渲染精炼参数估计。开发了硬件原型,在仿真和受控真实采集中验证了该方法能有效基于无纹理3D模型恢复物体位姿,并展示了对其他参数化场景的初步成功结果。同时通过实验探究了该成像方案的极限与能力。
原文摘要 · Abstract (English)
We aim to recover the geometry of 3D parametric scenes using very few depth measurements from low-cost, commercially available time-of-flight sensors. These sensors offer very low spatial resolution (i.e., a single pixel), but image a wide field-of-view per pixel and capture detailed time-of-flight data in the form of time-resolved photon counts. This time-of-flight data encodes rich scene information and thus enables recovery of simple scenes from sparse measurements. We investigate the feasibility of using a distributed set of few measurements (e.g., as few as 15 pixels) to recover the geometry of simple parametric scenes with a strong prior, such as estimating the 6D pose of a known object. To achieve this, we design a method that utilizes both feed-forward prediction to infer scene parameters, and differentiable rendering within an analysis-by-synthesis framework to refine the scene parameter estimate. We develop hardware prototypes and demonstrate that our method effectively recovers object pose given an untextured 3D model in both simulations and controlled real-world captures, and show promising initial results for other parametric scenes. We additionally conduct experiments to explore the limits and capabilities of our imaging solution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。