将多步骤文字指令拼成连贯视频演示,解决传统方法无法处理复杂流程的问题。
Stitch-a-Demo: Video Demonstrations from Multistep Descriptions
- 基于检索构建多步视频演示,整合不同来源片段确保内容准确
- 在真实教学视频上提升至最新水平,人类偏好测试胜率达29%以上
- 适合需要生成复杂操作流程视频的场景,如教程制作与智能助手
当前视觉生成方法通常仅处理单一句子描述(如标题或动作说明),并据此检索或生成对应视觉内容。然而,现有工作难以实现多步骤描述(如烹饪食谱或园艺指南)的视觉化展示,若孤立处理每一步,会导致视频不连贯。本文提出Stitch-a-Demo,一种基于检索的新方法,从多步骤文本中组装出既准确又连贯的视频演示。该方法通过大规模弱监督数据训练,包含多样化流程并引入困难负样本,以增强结果的正确性与整体一致性。在真实场景教学视频上验证,Stitch-a-Demo达到领先性能,相比基线提升最高达29%,并在人类偏好评估中显著胜出。
原文摘要 · Abstract (English)
When obtaining visual illustrations from text descriptions, today's methods take a description with a single text context - a caption, or an action description - and retrieve or generate the matching visual context. However, prior work does not permit visual illustration of multistep descriptions, e.g. a cooking recipe or a gardening instruction manual, and simply handling each step description in isolation would result in an incoherent demonstration. We propose Stitch-a-Demo, a novel retrieval-based method to assemble a video demonstration from a multistep description. The resulting video contains clips, possibly from different sources, that accurately reflect all the step descriptions, while being visually coherent. We formulate a training pipeline that creates large-scale weakly supervised data containing diverse procedures and injects hard negatives that promote both correctness and coherence. Validated on in-the-wild instructional videos, Stitch-a-Demo achieves state-of-the-art performance, with gains up to 29% as well as dramatic wins in a human preference study.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。