arXiv:2606.30608cs.CV2026-06

用对话式智能体从文本或图像零样本重建完整可动3D物体

UnfoldArt: Zero-Shot Recovery of Full Articulated 3D Objects from Text or Image

论文配图:UnfoldArt: Zero-Shot Recovery of Full Articulated 3D Objects from Text or Image
图 1 · 摘自论文原文
  • 通过高低层智能体分轮辩论,结合视觉语言模型推理语义与运动
  • 利用生成视频先验还原被遮挡的内部结构,实现动态视角下几何暴露
  • 适合需要零样本生成复杂可动3D对象的机器人、虚拟现实研究者

可动3D物体在具身AI、机器人和虚拟现实中至关重要,但仅凭稀疏观测恢复其结构与运动仍具挑战。现有方法受限于缺乏监督数据或缺少可靠先验,难以复现关节结构、隐藏几何与内部构造。我们提出首个基于辩论驱动的代理框架,可从文本或图像输入中零样本重建完整可动3D物体。高层代理借助视觉-语言与视频模型进行语义与运动推理,低层代理估计关节参数与交互点;两者通过两轮结构化辩论,先利用全局-局部不一致,再以自由生成的视频为锚点。同一视频先验在达成共识后,驱动各部件运动以暴露静态视图无法推断的遮挡内部与几何。该方法联合推断关节结构并重建完整3D可动物体,生成高保真几何、内部结构及运动一致性状态。

原文摘要 · Abstract (English)

Articulated 3D objects are essential for interactive environments in embodied AI, robotics, and virtual reality, but reconstructing their structure and motion from sparse observations remains challenging. Existing approaches remain largely constrained by lack of supervised data or lack the priors needed to reliably recover articulation, hidden geometry, and internal object structure. We present the first debate-driven agentic approach to articulated 3D object reconstruction from text or image inputs that both grounds articulation reasoning in concrete motion and exposes the occluded geometry revealed under articulation. High-level agents reason about object semantics and motion using knowledge from vision-language and video models, while low-level agents estimate articulation parameters and interaction points; together, they engage in a two-round structured debate that first exploits global--local disagreement and then grounds the agents in freely generated video. The same video prior, conditioned on the agreed articulation, then drives each part through its motion to expose occluded interiors and geometry that cannot be inferred from a single static view. By combining agentic reasoning with a video generative prior, our approach jointly infers articulation and reconstructs complete 3D articulated objects, producing high-fidelity geometry, internal structure, and motion-consistent states beyond directly observed surfaces.

3D重建可动物体生成模型智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。