用语言指令统一生成多种视角图像,支持五种输入形式。
MVLLaVA: An Intelligent Agent for Unified and Flexible Novel View Synthesis
- 融合多视角扩散模型与LLaVA,通过指令引导生成新视角。
- 在多种输入下表现稳定,支持单图、描述、视角变化等五类任务。
- 适合需要灵活生成视角的视觉应用开发者或研究者。
本文提出MVLLaVA,一个面向新视角合成任务的智能代理。它将多个多视角扩散模型与大语言多模态模型LLaVA结合,可高效处理多样化任务。该系统支持五种不同输入类型:单张图像、文本描述、特定视角方位变化等,并在语言指令引导下生成对应的新视角图像。通过精心设计的任务专用指令模板对LLaVA进行微调,使系统具备根据用户指令生成新视角图像的能力。实验验证了其在多样任务中的鲁棒性能与广泛适用性,展现出强大的灵活性和统一性。
原文摘要 · Abstract (English)
This paper introduces MVLLaVA, an intelligent agent designed for novel view synthesis tasks. MVLLaVA integrates multiple multi-view diffusion models with a large multimodal model, LLaVA, enabling it to handle a wide range of tasks efficiently. MVLLaVA represents a versatile and unified platform that adapts to diverse input types, including a single image, a descriptive caption, or a specific change in viewing azimuth, guided by language instructions for viewpoint generation. We carefully craft task-specific instruction templates, which are subsequently used to fine-tune LLaVA. As a result, MVLLaVA acquires the capability to generate novel view images based on user instructions, demonstrating its flexibility across diverse tasks. Experiments are conducted to validate the effectiveness of MVLLaVA, demonstrating its robust performance and versatility in tackling diverse novel view synthesis challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。