arXiv:2501.02527cs.CV2025-01

用视觉输入动态生成文本提示,提升图像生成质量。

Vision-Driven Prompt Optimization for Large Language Models in Multimodal Generative Tasks

  • 根据视觉输入实时生成优化提示词,驱动高质量图像合成。
  • 在COCO和Sketchy数据集上FID、LPIPS等指标显著优于现有方法。
  • 适合需要跨模态生成的AI研究者与开发者使用。

视觉生成仍是人工智能领域的前沿挑战,需实现视觉理解与生成能力的无缝融合。本文提出一种新框架——视觉驱动提示优化(VDPO),利用大语言模型从视觉输入中动态生成文本提示,指导高保真图像合成。VDPO结合视觉嵌入提示调优器、文本指令生成器与视觉生成模块,在COCO和Sketchy等基准测试中持续超越现有方法,显著提升FID、LPIPS及BLEU/CIDEr得分。额外分析表明,VDPO具备良好的可扩展性、鲁棒性与泛化能力,适用于领域内与领域外任务。人类评估进一步验证其在生成视觉美观且语义连贯输出方面的实际优势。

原文摘要 · Abstract (English)

Vision generation remains a challenging frontier in artificial intelligence, requiring seamless integration of visual understanding and generative capabilities. In this paper, we propose a novel framework, Vision-Driven Prompt Optimization (VDPO), that leverages Large Language Models (LLMs) to dynamically generate textual prompts from visual inputs, guiding high-fidelity image synthesis. VDPO combines a visual embedding prompt tuner, a textual instruction generator, and a vision generation module to achieve state-of-the-art performance in diverse vision generation tasks. Extensive experiments on benchmarks such as COCO and Sketchy demonstrate that VDPO consistently outperforms existing methods, achieving significant improvements in FID, LPIPS, and BLEU/CIDEr scores. Additional analyses reveal the scalability, robustness, and generalization capabilities of VDPO, making it a versatile solution for in-domain and out-of-domain tasks. Human evaluations further validate the practical superiority of VDPO in generating visually appealing and semantically coherent outputs.

多模态生成提示优化视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。