arXiv:2510.04536cs.GRcs.AI2025-10被引 3

用自然语言生成3D内容,通过AI自动操作设计工具提升效率。

3Dify: a Framework for Procedural 3D-CG Generation Assisted by LLMs Using MCP and RAG

  • 通过MCP和GUI自动化技术,让LLM直接控制3D软件
  • 用户选图反馈后,模型学习偏好并改进后续生成结果
  • 支持本地部署大模型,降低使用成本适合创意工作者

本文提出「3Dify」,一个基于大型语言模型(LLMs)的程序化3D计算机图形(3D-CG)生成框架。该框架使用户仅需自然语言指令即可生成3D-CG内容。3Dify构建于开源AI应用开发平台Dify,融合了模型上下文协议(MCP)和检索增强生成(RAG)等前沿技术。为支持3D-CG生成,3Dify通过MCP自动化各类数字内容创作(DCC)工具的操作;当DCC工具不支持MCP时,则采用计算机使用代理(CUA)方法自动化图形用户界面(GUI)操作。此外,为提升图像生成质量,3Dify允许用户从多个候选图像中选择偏好项,大模型据此学习变量模式,并应用于后续生成。同时,3Dify支持本地部署的LLMs,使用户可使用自研模型,利用自有计算资源减少外部API调用带来的时间与费用开销。

原文摘要 · Abstract (English)

This paper proposes "3Dify," a procedural 3D computer graphics (3D-CG) generation framework utilizing Large Language Models (LLMs). The framework enables users to generate 3D-CG content solely through natural language instructions. 3Dify is built upon Dify, an open-source platform for AI application development, and incorporates several state-of-the-art LLM-related technologies such as the Model Context Protocol (MCP) and Retrieval-Augmented Generation (RAG). For 3D-CG generation support, 3Dify automates the operation of various Digital Content Creation (DCC) tools via MCP. When DCC tools do not support MCP-based interaction, the framework employs the Computer-Using Agent (CUA) method to automate Graphical User Interface (GUI) operations. Moreover, to enhance image generation quality, 3Dify allows users to provide feedback by selecting preferred images from multiple candidates. The LLM then learns variable patterns from these selections and applies them to subsequent generations. Furthermore, 3Dify supports the integration of locally deployed LLMs, enabling users to utilize custom-developed models and to reduce both time and monetary costs associated with external API calls by leveraging their own computational resources.

3D生成大模型自动化AI创作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。