用概率模型提升低成本摄像头下人体定位精度。
Mean of Means: Human Localization with Calibration-free and Unconstrained Camera Settings (extended version)
- 将人体各点视为中心分布的观测值,大幅增加采样数
- 96%定位准确率在0.3米误差内,0.5米内接近100%
- 仅需两台640×480摄像头,成本低至10美元
精准的人体定位对元宇宙等应用至关重要。现有高精度方案依赖昂贵且需标记的硬件,而基于视觉的方法虽更便宜且无需标记,但现有立体视觉方法受限于刚性视角变换和多阶段SVD求解器中的误差传播,且需多台高分辨率相机并严格布设。为此,我们提出一种概率方法,将人体所有点视为以身体几何中心为均值的分布样本,显著提升采样效率,使每一点的采样数从数百增至数十亿。通过建模世界坐标与像素坐标均值之间的关系,利用中心极限定理保证正态性,从而促进学习。实验表明,在仅使用两台640×480分辨率网络摄像头、总成本仅10美元的条件下,人体定位在0.3米范围内准确率达96%,在0.5米范围内接近100%。
原文摘要 · Abstract (English)
Accurate human localization is crucial for various applications, especially in the Metaverse era. Existing high precision solutions rely on expensive, tag-dependent hardware, while vision-based methods offer a cheaper, tag-free alternative. However, current vision solutions based on stereo vision face limitations due to rigid perspective transformation principles and error propagation in multi-stage SVD solvers. These solutions also require multiple high-resolution cameras with strict setup constraints.To address these limitations, we propose a probabilistic approach that considers all points on the human body as observations generated by a distribution centered around the body's geometric center. This enables us to improve sampling significantly, increasing the number of samples for each point of interest from hundreds to billions. By modeling the relation between the means of the distributions of world coordinates and pixel coordinates, leveraging the Central Limit Theorem, we ensure normality and facilitate the learning process. Experimental results demonstrate human localization accuracy of 96\% within a 0.3$m$ range and nearly 100\% accuracy within a 0.5$m$ range, achieved at a low cost of only 10 USD using two web cameras with a resolution of 640$\times$480 pixels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。