arXiv:2508.08821cs.CV2025-08被引 2

仅用预训练多模态大模型生成带几何与部件标签的3D原型。

3DFroMLLM: 3D Prototype Generation only from Pretrained Multimodal LLMs

  • 通过设计-编码-视觉检查的智能体循环生成3D原型。
  • 生成图像用于分类预训练,性能比之前方法高15%。
  • 无需额外标注数据,可提升细粒度视觉语言模型55%准确率。

近期多模态大语言模型(MLLMs)在文本与图像联合表征学习方面表现出色,但其空间推理能力仍受限。本文提出3DFroMLLM,一种新框架,可直接从MLLMs生成包含几何结构和部件标签的3D物体原型。该流程采用智能体架构,包括设计师、编码器和视觉检查员,在迭代优化中完成生成。值得注意的是,该方法无需额外训练数据或详细用户指令。基于先前2D生成工作,我们证明由本框架生成的渲染图像可用于图像分类预训练,性能优于以往方法15%。作为真实应用场景,生成的原型可被用于微调CLIP以实现部件分割,无需依赖额外人工标注数据,即可使细粒度视觉语言模型准确率提升55%。

原文摘要 · Abstract (English)

Recent Multi-Modal Large Language Models (MLLMs) have demonstrated strong capabilities in learning joint representations from text and images. However, their spatial reasoning remains limited. We introduce 3DFroMLLM, a novel framework that enables the generation of 3D object prototypes directly from MLLMs, including geometry and part labels. Our pipeline is agentic, comprising a designer, coder, and visual inspector operating in a refinement loop. Notably, our approach requires no additional training data or detailed user instructions. Building on prior work in 2D generation, we demonstrate that rendered images produced by our framework can be effectively used for image classification pretraining tasks and outperforms previous methods by 15%. As a compelling real-world use case, we show that the generated prototypes can be leveraged to improve fine-grained vision-language models by using the rendered, part-labeled prototypes to fine-tune CLIP for part segmentation and achieving a 55% accuracy improvement without relying on any additional human-labeled data.

3D生成多模态智能体少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。