arXiv:2607.09365cs.RO2026-07

将视频生成的物体运动转为机器人可执行轨迹,提升成功率与可行性。

PhysV2A: Reachability-Gated and Semantic-Mask-Constrained Feasibility Completion for Video-to-Robot Manipulation

论文配图:PhysV2A: Reachability-Gated and Semantic-Mask-Constrained Feasibility Completion for Video-to-Robot Manipulation
图 1 · 摘自论文原文
  • 基于可达性门控和语义掩码约束,筛选可行抓取-轨迹组合。
  • 在4个桌面上操作任务中,成功率高于基准方法,失败率降低。
  • 适合需要高精度动作规划的机器人视觉操控场景。

基于视频的操作能从人类示范、生成视频或RGB-D观测中获取以物体为中心的运动先验,但这些先验通常与具体机器人无关,无法直接执行。本文提出PhysV2A,一种可达性门控与语义掩码约束的可行性补全框架,用于将视频导出的6D物体运动转换为机器人可执行的操作轨迹。核心思路是将抓取可行性视为轨迹条件化的而非局部的:每个由RGB-D生成的6自由度抓取候选,与恢复的物体运动刚性耦合,形成抓取条件化的工具中心点(TCP)轨迹假设。PhysV2A采用分层可达性门控选择,通过机器人本体运动学检查排除不可行的抓取-轨迹对,并根据下游执行适宜性对剩余候选排序。对选定的可达轨迹,利用视觉语言模型辅助和规则验证的S-Mask识别任务关键与可放松的笛卡尔分量,通过冗余优先优化与有界笛卡尔松弛实现语义掩码约束下的可操作性精炼。在四个桌面操作任务的真实机器人实验表明,PhysV2A相较于代表性视频先验和仅逆运动学基线方法,提升了任务成功率,减少了运动学可行性失败,并生成了更优条件的轨迹,其语义偏差保持有界。

原文摘要 · Abstract (English)

Video-based manipulation provides object-centric motion priors from human demonstrations, generated videos, or RGB-D observations, but such priors are typically embodiment-agnostic and cannot be directly executed by a specific robot. This paper presents \textbf{PhysV2A}, a reachability-gated and semantic-mask-constrained feasibility-completion framework for converting video-derived 6D object motion into robot-executable manipulation trajectories. The key idea is to treat grasp feasibility as trajectory-conditioned rather than local: each RGB-D-generated 6-DoF grasp candidate is rigidly coupled with the recovered object motion to form a grasp-conditioned TCP trajectory hypothesis. PhysV2A then performs hierarchical reachability-gated selection, where infeasible grasp--trajectory pairs are rejected by robot-centric kinematic checks and surviving candidates are ranked by downstream execution suitability. For the selected reachable trajectory, a VLM-assisted and rule-validated S-Mask identifies task-critical and relaxable Cartesian components, enabling semantic-mask-constrained manipulability refinement through redundancy-first optimization and bounded Cartesian relaxation. Real-robot experiments on four tabletop manipulation tasks show that PhysV2A improves task success over representative video-prior and IK-only baselines, reduces kinematic-feasibility failures, and produces better-conditioned trajectories with bounded semantic deviations.

机器人操控视频到动作运动规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。