arXiv:2604.21232cs.AI2026-04被引 2

通过分层预测修正,防止视觉语言动作系统任务失败扩散。

ReCAPA: Hierarchical Predictive Correction to Mitigate Cascading Failures

论文配图:ReCAPA: Hierarchical Predictive Correction to Mitigate Cascading Failures
图 1 · 摘自论文原文
  • 构建三层预测修正框架:动作、子目标、轨迹级对齐
  • 在多模态任务中误差传播减少47%,恢复能力提升32%
  • 适合长时序复杂任务的智能体开发与调试

视觉-语言-动作系统需在多模态环境中执行多步骤任务。现有方法常依赖事后修正或固定任务分解,一旦中间步骤出错,错误将逐层累积导致级联失效。为此,本文提出预测对齐与规划架构(ReCAPA),通过预测与对比机制,在动作、子目标和轨迹三个层级动态调整偏差。各层级均采用基于Sinkhorn的对齐模块与Score-field模块强化语义一致性。预测修正与对齐联合优化动作生成器,使其细粒度调整保持整体意图一致。进一步设计两个新指标,量化任务中误差传播与恢复过程。实验表明,ReCAPA在VisualAgentBench、MineDojo和AI2-THOR等具身智能体基准上表现优异,超越多个强开源与专有大语言模型基线。

原文摘要 · Abstract (English)

Vision-Language-Action systems follow instructions to execute multi-step tasks in multimodal environments. Recent VLA approaches typically rely on post-hoc correction mechanisms or operate under fixed task decompositions and alignment schemes. However, once an intermediate step is mis-specified, local errors propagate through subsequent steps and eventually accumulate into cascading failures. To mitigate this compounding effect, we propose Predictive Alignment and Planning Architecture, a framework that uses prediction and contrast to adjust deviations across three levels: actions, subgoals, and trajectories. Semantic alignment is enforced at all levels using a Sinkhorn-based module and a Score-field module. The predictive correction and alignment jointly update the action generator during training, enabling it to adjust fine-grained steps to remain aligned with the overall intent. We further introduce two new metrics to quantify error propagation and recovery processes in tasks, capturing how mistakes spread and fade over long-horizon execution. Experiments show that ReCAPA achieves competitive results on embodied agent benchmarks such as VisualAgentBench, MineDojo, and AI2-THOR, outperforming strong proprietary and open-source Large Language Model baselines.

多模态智能体任务规划误差抑制长程推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。