arXiv:2507.03313cs.CVcs.AI2025-07

用文字风格生成匹配作者气质的图像,让写作特色变视觉形象。

Personalized Image Generation from an Author Writing Style

  • 通过作者风格表输入LLM,生成三组图文提示词
  • 人类评估显示风格匹配度达4.08/5,图像具中等区分度
  • 适合创意辅助与跨模态理解场景,能捕捉情绪氛围

将细腻的文字化作者写作风格转化为引人入胜的视觉表现,是生成式AI中的新挑战。本文提出一个端到端流程:以作者写作特征的结构化摘要(Author Writing Sheets, AWS)为输入,通过大语言模型(Claude 3.7 Sonnet)生成三个描述性文生图提示词,再由扩散模型(Stable Diffusion 3.5 Medium)渲染成图像。我们在来自Reddit的49位作者风格上进行了评估,由人工评价者判断生成图像与文本风格的契合度及视觉独特性。结果显示,生成图像与作者文本特征在感知上具有良好一致性(平均风格匹配度:4.08/5),图像被评定为中等程度独特。定性分析进一步表明该流程能有效捕捉情绪与氛围,但在表达高度抽象的叙事元素方面仍存挑战。本研究提出了视觉作者风格个性化的创新方法,并提供了初步实证验证,为创意辅助与跨模态理解开辟了新路径。

原文摘要 · Abstract (English)

Translating nuanced, textually-defined authorial writing styles into compelling visual representations presents a novel challenge in generative AI. This paper introduces a pipeline that leverages Author Writing Sheets (AWS) - structured summaries of an author's literary characteristics - as input to a Large Language Model (LLM, Claude 3.7 Sonnet). The LLM interprets the AWS to generate three distinct, descriptive text-to-image prompts, which are then rendered by a diffusion model (Stable Diffusion 3.5 Medium). We evaluated our approach using 49 author styles from Reddit data, with human evaluators assessing the stylistic match and visual distinctiveness of the generated images. Results indicate a good perceived alignment between the generated visuals and the textual authorial profiles (mean style match: $4.08/5$), with images rated as moderately distinctive. Qualitative analysis further highlighted the pipeline's ability to capture mood and atmosphere, while also identifying challenges in representing highly abstract narrative elements. This work contributes a novel end-to-end methodology for visual authorial style personalization and provides an initial empirical validation, opening avenues for applications in creative assistance and cross-modal understanding.

风格迁移文生图个性化生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。