用3D合成数据提升多模态模型对相机与物体关系的理解能力。
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation
- 通过3D资产生成逼真图像,精确控制相机与物体关系。
- 构建24万条带标注的视觉问答数据集,准确率提升33.4%。
- 适合研究视觉理解、3D生成和多模态模型的开发者使用。
多模态大语言模型在捕捉相机-物体关系(如物体朝向、相机视角、拍摄镜头)方面表现不佳,主要因训练数据中此类关系多样性不足。为此,我们提出一种合成生成流程,构建大规模3D视觉指令数据集。该框架以3D资产为输入,结合渲染与基于扩散的图像生成模型,生成保持精确相机-物体关系的逼真图像;同时利用大语言模型生成文本提示,用于指导视觉指令微调与图像生成控制。我们构建了包含24万条视觉问答的Ultimate3D数据集及对应基准测试。在该数据集上微调的多模态模型显著优于商用模型,在相机-物体关系识别任务上平均准确率提升33.4%。相关代码、数据集与基准将推动多模态大模型广泛应用。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) struggle with accurately capturing camera-object relations, especially for object orientation, camera viewpoint, and camera shots. This stems from the fact that existing MLLMs are trained on images with limited diverse camera-object relations and corresponding textual descriptions. To address this, we propose a synthetic generation pipeline to create large-scale 3D visual instruction datasets. Our framework takes 3D assets as input and uses rendering and diffusion-based image generation models to create photorealistic images preserving precise camera-object relations. Additionally, large language models (LLMs) are used to generate text prompts for guiding visual instruction tuning and controlling image generation. We create Ultimate3D, a dataset of 240K VQAs with precise camera-object annotations, and corresponding benchmark. MLLMs fine-tuned on our proposed dataset outperform commercial models by a large margin, achieving an average accuracy improvement of 33.4% on camera-object relation recognition tasks. Our code, dataset, and benchmark will contribute to broad MLLM applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。