Iron让虚拟助手更懂指令,自动从失败中学习,减少标注依赖。
Iron: Intent-Aligned and Retrospective Dual Learning Framework for Enhancing Generalist Virtual Agents

- 用分步循环一致奖励,精准对齐动作与高阶意图
- 复用失败轨迹训练,提升学习效率和任务多样性
- 在未见过的网页任务上性能提升25.06%,适合复杂任务场景
实现能在多种数字环境中自动执行任务的虚拟代理仍是具身AI的核心挑战。尽管多模态大语言模型(MLLM)提升了视觉感知与推理能力,其作为代理部署仍面临三大难题:高昂的数据标注成本、动作与意图的不精确对齐,以及因丢弃失败轨迹导致的探索效率低下。为此,我们提出Iron框架,一种意图对齐、自我优化且标注高效的GUI代理训练方法。Iron采用新颖的双重学习策略,利用分步循环一致(SCC)奖励,实现低层动作与高层意图之间的细粒度对齐,从而提升指令定位与意图理解能力。同时,引入事后重演机制,将失败轨迹重新用于训练,提高学习效率与任务多样性。大量实验表明,Iron训练的通用代理在跨环境、跨设备任务中持续表现更优,性能超越使用三倍数据训练的模型。在未见过的网页任务上相对提升达25.06%,复杂任务中进一步取得增益,验证了构建更强大虚拟代理的可行性。
原文摘要 · Abstract (English)
Achieving virtual agents capable of automating tasks across diverse digital environments remains a pivotal challenge in Embodied AI. While Multimodal Large Language Models (MLLMs) offer enhanced visual perception and reasoning, their agentic deployment faces three challenges: costly data annotation, imprecise action-intent alignment, and inefficient exploration from discarded failed trajectories. To address these, we introduce Iron, an intent-aligned, self-improved, and annotation-efficient framework for training GUI agents. Iron employs a novel dual learning strategy that utilizes a stepwise cycle-consistent (SCC) reward to achieve fine-grained alignment between low-level actions and high-level intents, thereby improving instruction grounding and intent understanding. Concurrently, Iron introduces a hindsight reproduction mechanism to repurpose failed trajectories for training, improving both learning efficiency and task diversity. Extensive experiments demonstrate that Iron-trained generalist agents consistently improve performance on cross-environment and cross-device tasks, outperforming models trained with three times more data. Iron also achieves a substantial 25.06% relative improvement on unseen web tasks, with further gains observed on inherently complex tasks, demonstrating the feasibility of building more capable virtual agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。