用伪标签集成提升视频语言模型生成流程操作指南能力
In-Context Ensemble Learning from Pseudo Labels Improves Video-Language Models for Low-Level Workflow Understanding
- 通过上下文集成学习聚合多种可能的步骤路径伪标签
- 在零样本条件下显著提升流程步骤的时间准确性
- 适合需要自动化生成软件操作流程文档的研究与应用
标准操作流程(SOP)是业务软件工作流的低层级、分步书面指导。自动创建SOP是实现端到端软件流程自动化的关键步骤。手动编写SOP耗时费力。近期大型视频-语言模型可通过分析人类示范视频,具备自动化生成SOP的潜力。然而,现有大模型在零样本场景下生成SOP仍存在挑战。本文首先探索了视频-语言模型在SOP生成中的上下文学习能力。随后提出一种以探索为导向的策略——上下文集成学习,用于聚合多个可能的SOP路径的伪标签。该方法使模型在超出其上下文窗口限制的情况下仍能通过隐式一致性正则化持续学习。实验表明,上下文学习有助于提升模型生成更符合时间顺序的SOP;所提方法能持续增强模型在零样本条件下的SOP生成性能。
原文摘要 · Abstract (English)
A Standard Operating Procedure (SOP) defines a low-level, step-by-step written guide for a business software workflow. SOP generation is a crucial step towards automating end-to-end software workflows. Manually creating SOPs can be time-consuming. Recent advancements in large video-language models offer the potential for automating SOP generation by analyzing recordings of human demonstrations. However, current large video-language models face challenges with zero-shot SOP generation. In this work, we first explore in-context learning with video-language models for SOP generation. We then propose an exploration-focused strategy called In-Context Ensemble Learning, to aggregate pseudo labels of multiple possible paths of SOPs. The proposed in-context ensemble learning as well enables the models to learn beyond its context window limit with an implicit consistency regularisation. We report that in-context learning helps video-language models to generate more temporally accurate SOP, and the proposed in-context ensemble learning can consistently enhance the capabilities of the video-language models in SOP generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。