通过自纠错循环提升视觉语言导航模型的纠错能力
CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model
- 利用错误轨迹生成自纠错数据,形成持续训练闭环
- 在R2R-CE和RxR-CE上分别达到65.1%和69.3%成功率
- 适合需要高鲁棒性导航的机器人应用
现有视觉语言导航模型在执行指令时容易偏离正确路径,且缺乏有效纠错能力。为此,我们提出自纠错飞轮(Self-correction Flywheel)后训练范式,将训练集中的错误轨迹视为宝贵数据源而非缺陷。我们设计方法识别错误轨迹,并自动生成感知与动作层面的自纠错数据,作为模型持续训练的驱动力。重新评估模型时,新发现的错误轨迹再次触发自纠错循环,形成迭代飞轮。通过多轮迭代,我们逐步优化基于单目RGB的VLA导航模型CorrectNav。在R2R-CE和RxR-CE基准测试中,CorrectNav分别取得65.1%和69.3%的成功率,超越此前最优模型8.2%和16.4%。真实机器人测试表明,该模型具备出色的纠错、动态避障与长指令跟随能力。
原文摘要 · Abstract (English)
Existing vision-and-language navigation models often deviate from the correct trajectory when executing instructions. However, these models lack effective error correction capability, hindering their recovery from errors. To address this challenge, we propose Self-correction Flywheel, a novel post-training paradigm. Instead of considering the model's error trajectories on the training set as a drawback, our paradigm emphasizes their significance as a valuable data source. We have developed a method to identify deviations in these error trajectories and devised innovative techniques to automatically generate self-correction data for perception and action. These self-correction data serve as fuel to power the model's continued training. The brilliance of our paradigm is revealed when we re-evaluate the model on the training set, uncovering new error trajectories. At this time, the self-correction flywheel begins to spin. Through multiple flywheel iterations, we progressively enhance our monocular RGB-based VLA navigation model CorrectNav. Experiments on R2R-CE and RxR-CE benchmarks show CorrectNav achieves new state-of-the-art success rates of 65.1% and 69.3%, surpassing prior best VLA navigation models by 8.2% and 16.4%. Real robot tests in various indoor and outdoor environments demonstrate \method's superior capability of error correction, dynamic obstacle avoidance, and long instruction following.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。