arXiv:2508.02890cs.CVcs.CL2025-08

通过结构化信息提取提升视觉引导创作的准确性与创意性。

VisuCraft: Enhancing Large Vision-Language Models for Complex Visual-Guided Creative Content Generation via Structured Information Extraction

  • 从图像中提取细粒度视觉属性,生成结构化描述
  • 动态生成优化提示,使大模型更精准理解用户需求
  • 在故事和诗歌生成任务中显著提升创意与指令遵循度

本文提出VisuCraft,一种用于增强大型视觉语言模型(LVLMs)在复杂视觉引导创造性内容生成方面能力的新框架。现有LVLMs在生成长文本时常面临视觉保真度不足、创意匮乏及对细微用户指令响应不准确的问题。VisuCraft通过引入多模态结构化信息提取器(E)和动态提示生成模块(G)解决上述挑战。提取器将输入图像中的细粒度视觉属性转化为丰富结构化表示,动态提示模块则将其与用户指令结合,生成针对底层LVLM(如LLaVA、InstructBLIP)的高度优化提示。在自建的ImageStoryGen-500K数据集上,基于VisuGen指标(视觉定位、创意性、指令遵循度)评估显示,VisuCraft在故事生成和诗歌创作等任务中持续优于基线模型。结果表明其在创意性和指令遵循性方面有显著提升,验证了该框架在生成富有想象力、视觉相关且符合用户意图的长文本内容方面的有效性。本工作为LVLM在复杂创意AI应用中的潜力开辟了新路径。

原文摘要 · Abstract (English)

This paper introduces VisuCraft, a novel framework designed to significantly enhance the capabilities of Large Vision-Language Models (LVLMs) in complex visual-guided creative content generation. Existing LVLMs often exhibit limitations in maintaining high visual fidelity, genuine creativity, and precise adherence to nuanced user instructions when generating long-form texts. VisuCraft addresses these challenges by integrating a multimodal structured information extractor (E) and a dynamic prompt generation module (G). The extractor distills fine-grained visual attributes from input images into a rich, structured representation, which the dynamic prompt module then combines with user instructions to create highly optimized prompts for underlying LVLMs (e.g., LLaVA, InstructBLIP). Evaluated on the self-constructed ImageStoryGen-500K dataset using VisuGen Metrics (Visual Grounding, Creativity, and Instruction Adherence), VisuCraft consistently outperforms baseline LVLMs across tasks like story generation and poetry composition. Our results demonstrate remarkable improvements, particularly in creativity and instruction adherence, validating VisuCraft's effectiveness in producing imaginative, visually grounded, and user-aligned long-form creative text. This work unlocks new potential for LVLMs in sophisticated creative AI applications.

视觉语言模型创意生成结构化提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。