让机器人理解短期意图,避免动作冲突,提升操作稳定性。
IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation

- 用历史视觉信息编码短期意图,指导动作生成
- 在12个任务上显著提升执行稳定性与成功率
- 适合需要精细动作规划的机器人场景
机器人模仿数据常具多模态性:相似的视觉-语言观测可能对应不同动作序列,因示范者具有不同短期意图、任务阶段或近期上下文。现有帧条件型视觉-语言代理(VLA)策略仅依赖当前观测和指令推理每段动作,部分可观测下会在相邻重规划步骤中反复采样不同意图,导致动作段间冲突和执行不稳定。本文提出IntentVLA,一种历史条件型VLA框架,将近期视觉观测编码为紧凑的短期意图表征,并用于条件化动作段生成。我们还构建了AliasBench,一个基于RoboTwin2的12任务歧义感知基准,包含匹配的训练数据与评估环境,专门隔离短期观测混淆问题。在AliasBench、SimplerEnv、LIBERO和RoboCasa多个数据集上,IntentVLA均提升轨迹稳定性,超越强基线模型。
原文摘要 · Abstract (English)
Robot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human demonstrators act with different short-horizon intents, task phases, or recent context. Existing frame-conditioned VLA policies infer each chunk from the current observation and instruction alone, so under partial observability they may resample different intents across adjacent replanning steps, leading to inter-chunk conflict and unstable execution. We introduce IntentVLA, a history-conditioned VLA framework that encodes recent visual observations into a compact short-horizon intent representation and uses it to condition chunk generation. We further introduce AliasBench, a 12-task ambiguity-aware benchmark on RoboTwin2 with matched training data and evaluation environments that isolate short-horizon observation aliasing. Across AliasBench, SimplerEnv, LIBERO, and RoboCasa, IntentVLA improves rollout stability and outperforms strong VLA baselines
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。