arXiv:2508.19244cs.CV2025-08被引 1

无需训练,用文字就能控制3D模型的姿势。

Articulate3D: Zero-Shot Text-Driven 3D Object Posing

  • 用文本指令控制图像生成,保持结构一致
  • 通过关键点匹配实现多视角姿态优化
  • 适合零样本设计与自由文本控制场景

我们提出一种无需训练的方法 Articulate3D,通过语言指令操控3D资产的姿态。尽管视觉与语言模型已有进展,该任务仍具挑战性。方法分为两步:首先修改强大的图像生成器,根据输入图像和文本指令生成目标图像;其次通过多视角姿态优化将网格对齐至目标图像。我们引入自注意力重连机制(RSActrl),解耦图像生成模型中的结构与姿态,确保不同姿态下结构一致性。实验表明,可微渲染在关节优化中不可靠,因此改用关键点建立输入与目标图像间的对应关系。Articulate3D 在多种3D物体及自由文本提示下均表现良好,成功调整姿态同时保留原始网格身份。定量评估与用户对比研究显示,该方法在85%情况下更受青睐,优于现有方法。

原文摘要 · Abstract (English)

We propose a training-free method, Articulate3D, to pose a 3D asset through language control. Despite advances in vision and language models, this task remains surprisingly challenging. To achieve this goal, we decompose the problem into two steps. We modify a powerful image-generator to create target images conditioned on the input image and a text instruction. We then align the mesh to the target images through a multi-view pose optimisation step. In detail, we introduce a self-attention rewiring mechanism (RSActrl) that decouples the source structure from pose within an image generative model, allowing it to maintain a consistent structure across varying poses. We observed that differentiable rendering is an unreliable signal for articulation optimisation; instead, we use keypoints to establish correspondences between input and target images. The effectiveness of Articulate3D is demonstrated across a diverse range of 3D objects and free-form text prompts, successfully manipulating poses while maintaining the original identity of the mesh. Quantitative evaluations and a comparative user study, in which our method was preferred over 85\% of the time, confirm its superiority over existing approaches. Project page:https://odeb1.github.io/articulate3d_page_deb/

3D生成文本控制姿态优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。