仅用单目视频零样本分析3D物体运动,无需标注数据
MonoMobility: Zero-Shot 3D Mobility Analysis from Monocular Videos
- 结合深度估计与光流分析构建场景几何,初步解析运动部件
- 通过2D高斯点云优化实现旋转、平移等复杂运动的精确建模
- 适用于真实与模拟场景,适合无人监督的智能体环境理解
准确分析动态环境中物体的运动部件及其运动属性,对推动具身智能等关键领域至关重要。针对现有方法依赖密集多视角图像或精细部件级标注的问题,本文提出一种全新的零样本3D运动分析框架,仅需单目视频即可解析运动部件及运动属性,完全无需标注训练数据。该方法首先利用深度估计、光流分析和点云配准构建场景几何,粗略分析运动部件及其初始运动属性;随后采用2D高斯点云表示场景。在此基础上,设计端到端动态场景优化算法,专用于刚性连接物体,进一步精炼初始结果,确保系统可处理‘旋转’、‘平移’乃至复合运动(‘旋转+平移’),展现高灵活性与通用性。为验证方法鲁棒性与广泛适用性,我们构建了涵盖仿真与真实场景的综合性数据集。实验表明,本框架能在无标注条件下有效分析铰接物体运动,展现出在未来的具身智能应用中巨大潜力。
原文摘要 · Abstract (English)
Accurately analyzing the motion parts and their motion attributes in dynamic environments is crucial for advancing key areas such as embodied intelligence. Addressing the limitations of existing methods that rely on dense multi-view images or detailed part-level annotations, we propose an innovative framework that can analyze 3D mobility from monocular videos in a zero-shot manner. This framework can precisely parse motion parts and motion attributes only using a monocular video, completely eliminating the need for annotated training data. Specifically, our method first constructs the scene geometry and roughly analyzes the motion parts and their initial motion attributes combining depth estimation, optical flow analysis and point cloud registration method, then employs 2D Gaussian splatting for scene representation. Building on this, we introduce an end-to-end dynamic scene optimization algorithm specifically designed for articulated objects, refining the initial analysis results to ensure the system can handle 'rotation', 'translation', and even complex movements ('rotation+translation'), demonstrating high flexibility and versatility. To validate the robustness and wide applicability of our method, we created a comprehensive dataset comprising both simulated and real-world scenarios. Experimental results show that our framework can effectively analyze articulated object motions in an annotation-free manner, showcasing its significant potential in future embodied intelligence applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。