用文字生成流程图,让复杂操作更易懂。
$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion
- 拆解指令为目标和步骤,用语言结构指导图像生成
- 在三个数据集上生成的图与文本匹配度显著优于基线
- 适合教育、操作指引等需要图文联动的场景
如何有效传达程序性知识仍是自然语言处理中的难题,纯文本常难以表达复杂动作与空间关系。本文提出一种语言驱动框架,将程序性文本转化为连贯的视觉说明。通过将指令分解为目标陈述与顺序步骤,并基于这些语言元素条件化视觉生成,我们引入三项创新:(1) 基于成分句法分析器的文本编码机制,即使在长指令下也保持语义完整;(2) 成对话语连贯性模型,确保指令序列间的一致性;(3) 针对程序性语言-图像对齐设计的新评估协议。在三个指令数据集(HTStep、CaptainCook4D、WikiAll)上的实验表明,本方法在生成准确反映语言内容与顺序性的图像方面显著优于现有基线。该研究推动了程序性语言在视觉内容中的具身化,适用于教育、任务引导及多模态理解等领域。
原文摘要 · Abstract (English)
The effective communication of procedural knowledge remains a significant challenge in natural language processing (NLP), as purely textual instructions often fail to convey complex physical actions and spatial relationships. We address this limitation by proposing a language-driven framework that translates procedural text into coherent visual instructions. Our approach models the linguistic structure of instructional content by decomposing it into goal statements and sequential steps, then conditioning visual generation on these linguistic elements. We introduce three key innovations: (1) a constituency parser-based text encoding mechanism that preserves semantic completeness even with lengthy instructions, (2) a pairwise discourse coherence model that maintains consistency across instruction sequences, and (3) a novel evaluation protocol specifically designed for procedural language-to-image alignment. Our experiments across three instructional datasets (HTStep, CaptainCook4D, and WikiAll) demonstrate that our method significantly outperforms existing baselines in generating visuals that accurately reflect the linguistic content and sequential nature of instructions. This work contributes to the growing body of research on grounding procedural language in visual content, with applications spanning education, task guidance, and multimodal language understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。