arXiv:2507.08513cs.GRcs.CV2025-07被引 2

用3D合成数据提升多模态模型对相机与物体关系的理解能力。

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation

  • 通过3D资产生成逼真图像,精确控制相机与物体关系。
  • 构建24万条带标注的视觉问答数据集,准确率提升33.4%。
  • 适合研究视觉理解、3D生成和多模态模型的开发者使用。

多模态大语言模型在捕捉相机-物体关系(如物体朝向、相机视角、拍摄镜头)方面表现不佳,主要因训练数据中此类关系多样性不足。为此,我们提出一种合成生成流程,构建大规模3D视觉指令数据集。该框架以3D资产为输入,结合渲染与基于扩散的图像生成模型,生成保持精确相机-物体关系的逼真图像;同时利用大语言模型生成文本提示,用于指导视觉指令微调与图像生成控制。我们构建了包含24万条视觉问答的Ultimate3D数据集及对应基准测试。在该数据集上微调的多模态模型显著优于商用模型,在相机-物体关系识别任务上平均准确率提升33.4%。相关代码、数据集与基准将推动多模态大模型广泛应用。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) struggle with accurately capturing camera-object relations, especially for object orientation, camera viewpoint, and camera shots. This stems from the fact that existing MLLMs are trained on images with limited diverse camera-object relations and corresponding textual descriptions. To address this, we propose a synthetic generation pipeline to create large-scale 3D visual instruction datasets. Our framework takes 3D assets as input and uses rendering and diffusion-based image generation models to create photorealistic images preserving precise camera-object relations. Additionally, large language models (LLMs) are used to generate text prompts for guiding visual instruction tuning and controlling image generation. We create Ultimate3D, a dataset of 240K VQAs with precise camera-object annotations, and corresponding benchmark. MLLMs fine-tuned on our proposed dataset outperform commercial models by a large margin, achieving an average accuracy improvement of 33.4% on camera-object relation recognition tasks. Our code, dataset, and benchmark will contribute to broad MLLM applications.

多模态3D生成视觉问答指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。