arXiv:2511.01294cs.ROcs.CV2025-11被引 8

从图片或文字自动生成高自由度可动物体模型

Kinematify: Open-Vocabulary Synthesis of High-DoF Articulated Objects

  • 用搜索与优化结合方法推断复杂物体的运动结构
  • 在真实和合成数据上提升拓扑与参数精度
  • 适合机器人操作、物理模拟等需要可动模型的场景

对运动结构和可动部件的深度理解是机器人操控物体及建模自身结构的基础。这类理解依赖于可动物体模型,对物理仿真、运动规划和策略学习至关重要。然而,尤其是对于高自由度(DoF)物体,构建此类模型仍面临重大挑战。现有方法通常依赖运动序列或手工标注数据集的强假设,难以扩展。本文提出Kinematify,一个直接从任意RGB图像或文本描述中自动合成可动物体的框架。该方法解决两个核心问题:(i) 推断高自由度物体的运动拓扑结构;(ii) 从静态几何中估计关节参数。我们结合蒙特卡洛树搜索(MCTS)进行结构推理,以及基于几何的优化进行关节推理,生成物理一致且功能有效的描述。我们在合成与真实环境的多样化输入上评估了Kinematify,结果表明其在注册精度和运动拓扑准确性方面优于以往方法。

原文摘要 · Abstract (English)

A deep understanding of kinematic structures and movable components is essential for enabling robots to manipulate objects and model their own articulated forms. Such understanding is captured through articulated objects, which are essential for tasks such as physical simulation, motion planning, and policy learning. However, creating these models, particularly for objects with high degrees of freedom (DoF), remains a significant challenge. Existing methods typically rely on motion sequences or strong assumptions from hand-curated datasets, which hinders scalability. In this paper, we introduce Kinematify, an automated framework that synthesizes articulated objects directly from arbitrary RGB images or textual descriptions. Our method addresses two core challenges: (i) inferring kinematic topologies for high-DoF objects and (ii) estimating joint parameters from static geometry. To achieve this, we combine MCTS search for structural inference with geometry-driven optimization for joint reasoning, producing physically consistent and functionally valid descriptions. We evaluate Kinematify on diverse inputs from both synthetic and real-world environments, demonstrating improvements in registration and kinematic topology accuracy over prior work.

可动物体机器人生成模型结构推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。