让视频编辑像改文字一样简单,通过文本重写实现智能重拍。
Rewriting Video: Text-Driven Reauthoring of Video Footage
- 将视频逆向生成可编辑的文本提示,实现文本驱动重编
- 12位创作者实测发现新用例,如虚拟重拍与风格重塑
- 揭示人机协同中的连贯性与创作对齐难题,适合内容创作者
视频是强大的传播与叙事媒介,但重编现有画面仍具挑战性。即使是简单修改也需专业知识、时间和精心规划,限制了创作者对叙事的想象与构建。近年来生成式AI的发展提出新范式:视频编辑能否如文本重写般简便?为此,我们提出一项技术探针与用户研究,探索文本驱动的视频重编。方法包含两项技术贡献:(1) 一种生成式重建算法,将视频逆向还原为可编辑的文本提示;(2) 一个交互式工具Rewrite Kit,支持创作者操控这些提示。算法的技术评估揭示了关键的人机感知差异。对12位创作者的探针研究发现了虚拟重拍、合成连贯性与美学重制等新应用场景,同时凸显了连贯性、控制力与创作对齐方面的核心矛盾。本研究为文本驱动视频重编提供了实证洞见,推动未来协同创作视频工具的设计。
原文摘要 · Abstract (English)
Video is a powerful medium for communication and storytelling, yet reauthoring existing footage remains challenging. Even simple edits often demand expertise, time, and careful planning, constraining how creators envision and shape their narratives. Recent advances in generative AI suggest a new paradigm: what if editing a video were as straightforward as rewriting text? To investigate this, we present a tech probe and a study on text-driven video reauthoring. Our approach involves two technical contributions: (1) a generative reconstruction algorithm that reverse-engineers video into an editable text prompt, and (2) an interactive probe, Rewrite Kit, that allows creators to manipulate these prompts. A technical evaluation of the algorithm reveals a critical human-AI perceptual gap. A probe study with 12 creators surfaced novel use cases such as virtual reshooting, synthetic continuity, and aesthetic restyling. It also highlighted key tensions around coherence, control, and creative alignment in this new paradigm. Our work contributes empirical insights into the opportunities and challenges of text-driven video reauthoring, offering design implications for future co-creative video tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。