从一张静态图推断物体关节结构,无需多视角或先验标注
DailyArt: Discovering Articulation from Single Static Images via Latent Dynamics
- 先合成展开状态暴露运动线索,再对比推理关节参数
- 单图即可同时估计全部关节,无需模板或标注
- 可基于估计关节生成新姿态,适合机器人感知任务
关节物体对具身智能和世界建模至关重要,但从单张闭合状态图像推断其运动学仍具挑战,因关键运动线索常被遮挡。现有方法需多状态观测或依赖显式部件先验、检索等辅助输入,部分暴露待推断结构。本文提出DailyArt,将单图关节估计转化为合成引导的推理问题:先在相同视角下合成最大展开状态以暴露关节线索,再通过观测与合成状态间的差异估计完整关节参数。采用集合预测形式,可无须对象特异性模板、多视角输入或测试时显式部件标注,一次性恢复所有关节。以估计关节为条件,框架还可支持部件级新状态合成这一下游能力。大量实验表明,DailyArt在关节估计上表现优异,并能基于关节生成新姿态。项目页面见https://rangooo123.github.io/DaliyArt.github.io/
原文摘要 · Abstract (English)
Articulated objects are essential for embodied AI and world models, yet inferring their kinematics from a single closed-state image remains challenging because crucial motion cues are often occluded. Existing methods either require multi-state observations or rely on explicit part priors, retrieval, or other auxiliary inputs that partially expose the structure to be inferred. In this work, we present DailyArt, which formulates articulated joint estimation from a single static image as a synthesis-mediated reasoning problem. Instead of directly regressing joints from a heavily occluded observation, DailyArt first synthesizes a maximally articulated opened state under the same camera view to expose articulation cues, and then estimates the full set of joint parameters from the discrepancy between the observed and synthesized states. Using a set-prediction formulation, DailyArt recovers all joints simultaneously without requiring object-specific templates, multi-view inputs, or explicit part annotations at test time. Taking estimated joints as conditions, the framework further supports part-level novel state synthesis as a downstream capability. Extensive experiments show that DailyArt achieves strong performance in articulated joint estimation and supports part-level novel state synthesis conditioned on joints. Project page is available at https://rangooo123.github.io/DaliyArt.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。