arXiv:2506.21832cs.CV2025-06

让用户成为故事主角,一键生成包含自己形象的个性化图文故事。

TaleForge: Interactive Multimodal System for Personalized Story Creation

  • 用大模型和扩散模型融合用户面部与穿搭生成角色画像。
  • 支持实时预览,用户参与度显著提升,情感投入更强。
  • 适合想创作专属故事的普通用户或内容创作者。

讲故事是高度个人化且富有创造力的过程,但现有方法常将用户视为被动消费者,提供通用情节且个性化程度有限,削弱了参与感与沉浸感,尤其在涉及个人风格或外貌时更为明显。我们提出TaleForge,一个结合大语言模型(LLMs)与文本到图像扩散模型的个性化故事生成系统,可将用户的面部图像嵌入叙事与插图中。该系统包含三个相互关联的模块:故事生成模块利用大模型根据用户提示生成情节与角色描述;个性化图像生成模块将用户面部与服装选择融合生成角色插图;背景生成模块则创建融入个性化角色的场景背景。用户研究显示,当个体作为主角出现时,参与感和归属感显著增强。参与者称赞系统具备实时预览与直观控制,但也希望获得更精细的叙事编辑功能。TaleForge通过协调个性化文本与图像,推动多模态叙事发展,实现以用户为中心的沉浸式体验。

原文摘要 · Abstract (English)

Storytelling is a deeply personal and creative process, yet existing methods often treat users as passive consumers, offering generic plots with limited personalization. This undermines engagement and immersion, especially where individual style or appearance is crucial. We introduce TaleForge, a personalized story-generation system that integrates large language models (LLMs) and text-to-image diffusion to embed users' facial images within both narratives and illustrations. TaleForge features three interconnected modules: Story Generation, where LLMs create narratives and character descriptions from user prompts; Personalized Image Generation, merging users' faces and outfit choices into character illustrations; and Background Generation, creating scene backdrops that incorporate personalized characters. A user study demonstrated heightened engagement and ownership when individuals appeared as protagonists. Participants praised the system's real-time previews and intuitive controls, though they requested finer narrative editing tools. TaleForge advances multimodal storytelling by aligning personalized text and imagery to create immersive, user-centric experiences.

个性化生成多模态故事创作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。