提出首个无需训练即可在真实场景中准确估计单目场景流的方法。
Zero-Shot Monocular Scene Flow Estimation in the Wild
- 联合预测几何与运动,提升估计精度。
- 构建100万样本合成数据集缓解标注稀缺问题。
- 零样本泛化到DAVIS和RoboTAP真实视频,适合实际应用。
大型模型已在深度估计等低层视觉任务中展现出跨数据集的泛化能力,但场景流领域尚无此类通用模型。尽管场景流具有广泛应用潜力,现有预测模型泛化能力差,限制了实际使用。本文识别出三大关键挑战并提出相应解决方案:首先,提出联合估计几何与运动的方法以提升预测准确性;其次,设计一种数据配方,生成涵盖多样合成场景的100万标注训练样本;第三,评估多种场景流参数化方式,采用自然且高效的形式。所提模型在3D端点误差上优于现有方法及基于大规模模型的基线,且在未见过的真实视频(DAVIS)和机器人操作场景(RoboTAP)上实现零样本泛化。整体方法使场景流预测更具现实可行性。
原文摘要 · Abstract (English)
Large models have shown generalization across datasets for many low-level vision tasks, like depth estimation, but no such general models exist for scene flow. Even though scene flow has wide potential use, it is not used in practice because current predictive models do not generalize well. We identify three key challenges and propose solutions for each. First, we create a method that jointly estimates geometry and motion for accurate prediction. Second, we alleviate scene flow data scarcity with a data recipe that affords us 1M annotated training samples across diverse synthetic scenes. Third, we evaluate different parameterizations for scene flow prediction and adopt a natural and effective parameterization. Our resulting model outperforms existing methods as well as baselines built on large-scale models in terms of 3D end-point error, and shows zero-shot generalization to the casually captured videos from DAVIS and the robotic manipulation scenes from RoboTAP. Overall, our approach makes scene flow prediction more practical in-the-wild.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。