让用户像玩玩具一样操控角色符号,自动生成故事文本和视觉内容。
Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols
- 通过角色符号动作与文本共享语义空间,实现动作与文本相互引导生成。
- 相比GPT-4o,用户在故事生成质量和表达意图上表现更优。
- 适合创意写作、儿童教育及人机交互研究者使用。
我们提出Toyteller,一个基于AI的故事生成系统,用户可通过像玩玩具一样直接操作角色符号,同时生成故事文本和视觉内容。拟人化的符号动作能传递丰富细腻的社会互动信息;Toyteller利用这些动作(1)引导故事文本生成,(2)作为与文本配套的视觉输出形式。通过将动作与文本映射到共享语义空间,实现大语言模型与动作生成模型之间的跨模态翻译。技术评估显示,Toyteller优于竞争性基线GPT-4o。用户研究发现,玩具式操作有助于表达难以用语言描述的意图,但仅靠动作仍无法完全表达所有意图,提示需结合语言等其他模态。本文探讨了玩具式交互的设计空间,并对人-人工智能交互的技术与设计研究具有启示意义。
原文摘要 · Abstract (English)
We introduce Toyteller, an AI-powered storytelling system where users generate a mix of story text and visuals by directly manipulating character symbols like they are toy-playing. Anthropomorphized symbol motions can convey rich and nuanced social interactions; Toyteller leverages these motions (1) to let users steer story text generation and (2) as a visual output format that accompanies story text. We enabled motion-steered text generation and text-steered motion generation by mapping motions and text onto a shared semantic space so that large language models and motion generation models can use it as a translational layer. Technical evaluations showed that Toyteller outperforms a competitive baseline, GPT-4o. Our user study identified that toy-playing helps express intentions difficult to verbalize. However, only motions could not express all user intentions, suggesting combining it with other modalities like language. We discuss the design space of toy-playing interactions and implications for technical HCI research on human-AI interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。