arXiv:2603.09072cs.HCcs.AI2026-03

用写文章方式直接生成视频,让创作更自然。

A Text-Native Interface for Generative Video Authoring

  • 以文本为核心,边写边构建视频元素与镜头
  • 用户可一键生成包含画面与音效的完整视频内容
  • 适合无视频经验者快速上手创作故事

人人都会写文字,但视频创作却需掌握复杂工具。本文提出 Doki,一种以文本为核心的生成式视频创作界面,将视频制作流程融入自然的写作习惯。在 Doki 中,用户通过单一文档定义素材、规划场景、设计镜头、调整剪辑并添加音频,所有操作均以文本为第一交互形式。我们阐述了这一文本优先的设计原则,并通过多个实例展示其能力。为评估实际应用效果,我们开展为期一周的部署研究,参与者涵盖不同视频制作经验者。该工作推动生成式视频界面的根本变革,提供了一种强大且易用的新方式来讲述视觉故事。

原文摘要 · Abstract (English)

Everyone can write their stories in freeform text format -- it's something we all learn in school. Yet storytelling via video requires one to learn specialized and complicated tools. In this paper, we introduce Doki, a text-native interface for generative video authoring, aligning video creation with the natural process of text writing. In Doki, writing text is the primary interaction: within a single document, users define assets, structure scenes, create shots, refine edits, and add audio. We articulate the design principles of this text-first approach and demonstrate Doki's capabilities through a series of examples. To evaluate its real-world use, we conducted a week-long deployment study with participants of varying expertise in video authoring. This work contributes a fundamental shift in generative video interfaces, demonstrating a powerful and accessible new way to craft visual stories.

视频生成文本接口创作工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。