arXiv:2503.05226cs.ROcs.AI2025-03被引 1

通过分层反馈信号提升机器人在高不确定性下的决策鲁棒性

Reward-Centered ReST-MCTS: A Robust Decision-Making Framework for Robotic Manipulation in High Uncertainty Environments

  • 将中间反馈拆解为规则、启发式、神经和估值四通道,聚焦任务上下文
  • 在相同模型基础上,无保护时0/10成功,有保护时9/10成功
  • 适合高不确定性环境下对已有视觉语言动作策略的可靠性验证

蒙特卡洛树搜索因其无需可微策略即可通过模拟优化动作选择,被广泛应用于机器人操作。但在不确定环境中,稀疏的终局奖励和噪声状态转移会使浅层搜索变得脆弱:许多候选分支直到后期回溯才可区分,且有限的仿真预算加剧了这种模糊性。本文提出奖励中心型ReST-MCTS(Reward-Centered ReST-MCTS),将中间反馈分解为规则、启发式、可选神经与估值估计四通道,基于匹配的任务上下文中心化处理信号,并用于引导或修复搜索过程,同时保留终局任务评估。主要证据分层设计:局部任务与匹配的ManiSkill诊断隔离奖励中心机制与消融实验;在相同基础模型下,通过匹配的选项级ManiSkill测试在原始姿态偏移、观测噪声及动作基元失效下验证鲁棒性,不宣称标准基准领先;官方同架构OpenVLA-OFT/LIBERO桥接实验验证动作修复能力。在相同模型下,开放复现达到10/10的LIBERO-Spatial成功率,无论是否启用RCRM-Guard;单套方案在十次动作通道压力测试中,未防护时0/10成功,防护后9/10成功。额外报告观察噪声、语言扰动与视觉干扰探测结果,作为覆盖范围与负向对照,非性能优势证明。结论限定:该方法是同架构高不确定性操作场景下的可解释性运行时验证器,非替代性视觉语言动作策略,亦不主张广泛基准优势。

原文摘要 · Abstract (English)

Monte Carlo tree search is attractive for robotic manipulation because it can improve action selection through simulation without requiring a fully differentiable policy. In uncertain domains, however, sparse terminal rewards and noisy transitions can make shallow search brittle: many candidate branches remain indistinguishable until late rollouts, and small simulation budgets amplify this ambiguity. This paper presents Reward-Centered ReST-MCTS, a decision-making framework that decomposes intermediate feedback into rule, heuristic, optional neural, and value-estimation channels, centers the resulting process signal against matched task contexts, and uses it to bias or repair search while preserving terminal-task evaluation. The primary evidence is intentionally tiered. Local tasks and matched ManiSkill diagnostics isolate reward-center mechanisms and ablations; matched option-level ManiSkill sweeps test robustness under primitive failure, observation noise, and initial-pose shifts while not claiming standard benchmark superiority; and an official same-backbone OpenVLA-OFT/LIBERO bridge tests bounded VLA action repair. The OpenVLA-OFT clean reproduction reaches 10/10 LIBERO-Spatial successes both with and without RCRM-Guard. A single-suite same-backbone action-channel stress artifact over ten paired LIBERO-Spatial action-channel stress episodes records 0/10 unguarded successes and 9/10 guarded successes. Additional observation-noise, language-perturbation, and visual-distractor probes are reported as coverage and negative-result context rather than superiority evidence. The resulting claim is bounded: Reward-Centered ReST-MCTS is an inspectable test-time verifier for same-backbone high-uncertainty manipulation, not a replacement VLA policy or a broad standard-benchmark superiority claim.

机器人操作强化学习决策框架不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。