arXiv:2511.03227cs.HCcs.AI2025-11中稿 · NeurIPS被引 4

用节点图编辑方式,实现多模态内容的可控创作。

Node-Based Editing for Multimodal Generation of Text, Audio, Image, and Video

  • 将故事建模为可编辑的节点图,支持多模态内容融合。
  • 实测支持自动生成故事大纲,用户可逐节点迭代优化。
  • 适合需要结构化创作的影视、游戏与交互叙事设计者。

我们提出一种基于节点的故事创作系统,用于多模态内容生成。系统将故事表示为可扩展、可编辑的节点图,通过直接修改或自然语言提示进行迭代优化。每个节点可集成文本、图像、音频和视频,支持创作者构建多模态叙事。任务选择代理负责调度专用生成任务,包括故事生成、节点结构推理、节点布局格式化和上下文生成。界面支持单个节点的精准编辑、自动分支生成并行剧情线,以及基于节点的迭代优化。实验结果表明,该方法能有效控制叙事结构,并实现文本、图像、音频和视频的渐进式生成。我们报告了自动故事大纲生成的量化结果及编辑流程的定性观察。最后讨论了当前在长篇叙事扩展性和多节点一致性方面的局限性,并提出未来向人机协同、以用户为中心的创意AI工具演进的方向。

原文摘要 · Abstract (English)

We present a node-based storytelling system for multimodal content generation. The system represents stories as graphs of nodes that can be expanded, edited, and iteratively refined through direct user edits and natural-language prompts. Each node can integrate text, images, audio, and video, allowing creators to compose multimodal narratives. A task selection agent routes between specialized generative tasks that handle story generation, node structure reasoning, node diagram formatting, and context generation. The interface supports targeted editing of individual nodes, automatic branching for parallel storylines, and node-based iterative refinement. Our results demonstrate that node-based editing supports control over narrative structure and iterative generation of text, images, audio, and video. We report quantitative outcomes on automatic story outline generation and qualitative observations of editing workflows. Finally, we discuss current limitations such as scalability to longer narratives and consistency across multiple nodes, and outline future work toward human-in-the-loop and user-centered creative AI tools.

多模态生成节点编辑叙事系统创意AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。