让机器人执行任务时实时判断动作是否可行并自我纠错,提升长周期操作可靠性。
PhysReflect-VLA: Physical Feasibility and Self-Reflective Regulation for Reliable Vision-Language-Action Policies

- 引入可行性评估与自我反思模块,在执行中闭环验证动作合理性。
- 在真实复杂任务中平均成功率提升5.4%,显著改善阶段稳定性和整体成功度。
- 适合需要高可靠性的长周期机器人操作场景,如工业装配与家庭服务。
长周期机器人操作对物理不可行动作、接触干扰及缺乏在线纠错极为敏感。尽管视觉-语言-动作(VLA)模型通过多模态学习具备强任务理解能力,但通常以前馈方式生成动作,未显式检查物理可行性或诊断执行错误。本文提出PhysReflect-VLA,一种可即插即用的执行期可靠性框架,通过闭环控制流程为VLA策略增加物理可行性评估与结构化自我反思能力。其中,可行性算子评估候选动作是否引发动态一致的状态转移;动作解释算子验证转移一致性;基于大语言模型的反思模块分析状态偏差,生成后续动作的修正指导。采用两阶段训练流程稳定可行性建模,并将反思集成至控制环路。在多阶段、高接触的真实世界操作任务中,实验表明该方法相较代表性VLA基线平均提升5.4%的任务成功率,且消融实验显示可行性检查与反思纠错均有助于增强执行鲁棒性。结果凸显了嵌入物理一致性验证与在线自我反思对可靠长周期机器人操作的重要性。
原文摘要 · Abstract (English)
Long-horizon robotic manipulation is highly sensitive to physically infeasible transitions, contact-induced disturbances, and the lack of effective self-correction during execution. Although Vision-Language-Action (VLA) models provide strong task grounding through multimodal learning, they typically generate actions in a feed-forward manner without explicitly checking physical feasibility or diagnosing execution errors online. We present PhysReflect-VLA, a plug-and-play execution-time reliability framework that augments VLA policies with physical feasibility evaluation and structured self-reflection in a closed-loop control pipeline. A Feasibility Operator evaluates whether candidate actions induce dynamically consistent state transitions; an Action Explanation Operator verifies transition coherence; and an LLM-based Reflection Module analyzes state discrepancies to generate corrective guidance for subsequent actions. A two-stage training procedure stabilizes feasibility modeling and integrates reflection into the control loop. Experiments on multi-stage, contact-rich real-world manipulation tasks show consistent improvements in stage-wise stability and overall task success compared with representative VLA baselines with an average gain of 5.4\%. Ablation results further indicate that feasibility checking and reflection-based correction both contribute to improved execution robustness. These results highlight the importance of embedding physical consistency checks and online self-reflection for reliable long-horizon robotic manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。