用图结构知识库让智能体自动执行复杂流程,适应变化环境。
Adaptive Multimodal Agents-Based Framework for Automatic Workflow Execution

- 分两阶段:离线构建流程拓扑图,线上基于图动态导航
- 在少量训练数据下仍保持高可靠性和语义理解能力
- 适合需要自适应执行的自动化工作流场景
现代信息系统需要能够自主导航复杂工作流的智能体,但现有方法在从结构化元数据解析过渡到通用环境感知时存在困难。尽管多模态大模型(MLLMs)使智能体能直接与图形界面交互,但现有方法通常将任务序列视为离散、线性的片段,导致无法捕捉底层状态转移拓扑,限制了其在新场景或非平稳环境中的表现。为此,我们提出一种基于自适应多模态智能体的新型框架,通过双阶段流程实现自动工作流执行:首先,在离线发现阶段,系统从碎片化执行日志中自适应构建拓扑知识库;推理时,智能体在固定预建图上使用自适应检索增强生成(RAG),结合闭环协作验证机制,实现动态自我修正与导航。该图结构方法显著提升任务分解与自适应导航性能。我们在真实场景中验证了该框架,证明其在有限训练数据下仍具备高可靠性与语义感知能力。
原文摘要 · Abstract (English)
Modern information systems require autonomous agents capable of navigating complex workflows, yet current methodologies often struggle with the transition from structured metadata parsing to general environmental perception. While the integration of MLLMs has enabled agents to interact directly with GUIs, existing approaches typically treat task sequences as discrete, linear episodes. This fragmentation prevents agents from capturing the underlying transition topology, limiting their effectiveness in novel or non-stationary scenarios. To address this, we propose a novel multimodal multi-agent framework that achieves automatic workflow execution through a distinct two-phase pipeline. First, during an offline discovery phase, the architecture adaptively constructs a topological knowledge base from fragmented execution logs. During inference, agents leverage Adaptive Retrieval-Augmented Generation (RAG) over this fixed, pre-established graph, coupled with a closed-loop collaborative verification protocol to dynamically self-correct and navigate. This graph-based approach facilitates superior task decomposition and adaptive navigation performance. We validate our framework in a real-world context, demonstrating its ability to maintain high reliability and semantic awareness even with limited training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。