用AI指令自动生成2D/3D/4D高质量数据,解决真实数据难收集问题
Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis
- 通过多模态大模型自动采集资产并生成3D布局
- 多视角场景优化与时间连贯帧生成提升数据质量
- 适合需要大规模合成数据的生成式AI研究者
随着AI生成内容(AIGC)需求增长,高质量、多样化且可扩展的数据愈发关键。然而,大规模真实世界数据的采集仍成本高昂且耗时,制约下游应用发展。尽管部分工作尝试通过渲染生成特定任务数据,多数方法仍依赖人工场景构建,限制了可扩展性与准确性。为此,我们提出Follow-Your-Instruction,一种由多模态大语言模型(MLLM)驱动的框架,用于自动合成高质量2D、3D和4D数据。该框架首先通过MLLM-Collector从多模态输入中收集资产及其描述;随后构建3D布局,并利用视觉-语言模型(VLM)在多视角场景中通过MLLM-Generator与MLLM-Optimizer进行语义精炼;最后,MLLM-Planner生成时间上连贯的未来帧。我们在2D、3D和4D生成任务上进行全面评估,结果表明合成数据显著提升现有基线模型性能,验证了Follow-Your-Instruction作为可扩展、高效数据引擎在生成智能中的潜力。
原文摘要 · Abstract (English)
With the growing demands of AI-generated content (AIGC), the need for high-quality, diverse, and scalable data has become increasingly crucial. However, collecting large-scale real-world data remains costly and time-consuming, hindering the development of downstream applications. While some works attempt to collect task-specific data via a rendering process, most approaches still rely on manual scene construction, limiting their scalability and accuracy. To address these challenges, we propose Follow-Your-Instruction, a Multimodal Large Language Model (MLLM)-driven framework for automatically synthesizing high-quality 2D, 3D, and 4D data. Our \textbf{Follow-Your-Instruction} first collects assets and their associated descriptions through multimodal inputs using the MLLM-Collector. Then it constructs 3D layouts, and leverages Vision-Language Models (VLMs) for semantic refinement through multi-view scenes with the MLLM-Generator and MLLM-Optimizer, respectively. Finally, it uses MLLM-Planner to generate temporally coherent future frames. We evaluate the quality of the generated data through comprehensive experiments on the 2D, 3D, and 4D generative tasks. The results show that our synthetic data significantly boosts the performance of existing baseline models, demonstrating Follow-Your-Instruction's potential as a scalable and effective data engine for generative intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。