arXiv:2601.14207cs.GRcs.CV2026-01被引 1

用文本提示零样本对齐两个3D物体,兼顾语义与物理合理性。

Copy-Trasform-Paste: Zero-Shot Object-Object Alignment Guided by Vision-Language and Geometric Constraints

  • 通过可微渲染结合CLIP梯度,实时优化物体位置姿态。
  • 在自建数据集上优于所有基线,实现语义准确且不穿透的对齐。
  • 适合需要快速生成合理3D场景的内容创作者使用。

我们研究基于文本提示描述空间关系的零样本3D物体对齐问题,这是内容创作与场景搭建的关键能力。现有方法多依赖几何对齐或利用预训练2D扩散模型建模语言-物体空间关系。本文提出直接在测试时优化相对位姿,通过可微渲染和CLIP驱动的梯度更新平移、旋转与等比缩放,无需训练新模型。框架引入几何感知目标:软ICP项促进表面贴合,穿透损失防止交叠,并采用分阶段策略强化接触约束,配合相机控制聚焦优化区域。为支持评估,构建包含多样类别与关系的基准数据集。实验表明,该方法在多个指标上优于现有基线,生成结果兼具语义一致性与物理合理性。

原文摘要 · Abstract (English)

We study zero-shot 3D alignment of two given meshes, using a text prompt describing their spatial relation -- an essential capability for content creation and scene assembly. Earlier approaches primarily rely on geometric alignment procedures, while recent work leverages pretrained 2D diffusion models to model language-conditioned object-object spatial relationships. In contrast, we directly optimize the relative pose at test time, updating translation, rotation, and isotropic scale with CLIP-driven gradients via a differentiable renderer, without training a new model. Our framework augments language supervision with geometry-aware objectives: a variant of soft-Iterative Closest Point (ICP) term to encourage surface attachment and a penetration loss to discourage interpenetration. A phased schedule strengthens contact constraints over time, and camera control concentrates the optimization on the interaction region. To enable evaluation, we curate a benchmark containing diverse categories and relations, and compare against baselines. Our method outperforms all alternatives, yielding semantically faithful and physically plausible alignments.

3D对齐视觉语言生成式几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。