arXiv:2609.02653cs.RO2026-09

让机器人像人一样理解模糊指令,持续完成复杂操作

HINT: Human-Intent Inception for Long-Horizon Robot Manipulation

论文配图:HINT: Human-Intent Inception for Long-Horizon Robot Manipulation
图 1 · 摘自论文原文
  • 仅在操作模式切换时进行语义推理,其余时间依赖视觉跟踪
  • 在三个长程任务中提升成功率,且不增加基础模型参数
  • 适合需要高可靠性、低延迟的机器人长程操作场景

人类能在简单指令下完成复杂操作,同时根据不断变化的视觉信息持续调整。然而,当前视觉-语言动作(VLA)模型及其他动作策略在密集、动态的视觉输入和稀疏语言引导下难以实现这种高层智能行为。视觉相关性常主导语义意图,导致动作跟随视觉捷径而非人类目标。我们提出HINT(Human-INTent INcepTion),一个受人类操作原理启发的代理框架:语义意图仅在操作模式转换时稀疏变化,而连续控制主要依赖物体与手之间的动态关系。HINT仅在模式转换时调用语义推理以解决当前子任务和目标,随后通过多视角定位和视觉跟踪维持该承诺。我们探索两种视觉接口——图像空间语义高亮和注意力优先注入——向动作策略传递追踪到的意图,无需向基础动作模型引入额外可训练参数。在三个长程任务及分布外变体上的实验表明,HINT显著提升了意图理解、任务进展和端到端成功率,同时保持低延迟控制。

原文摘要 · Abstract (English)

Humans can perform complex manipulations given a simple intent through an overall instruction, while continuously adapting to evolving visual observations. However, current vision-language action (VLA) models and other action policies struggle to realize this high-level intelligent behavior under dense, evolving visual inputs and sparse language guidance. Visual correlations can then dominate semantic intent, leading actions to follow visual shortcuts rather than human goals. We present HINT (Human-INTent INcepTion), an agentic framework inspired by the human manipulation principles: semantic intent changes sparsely at manipulation-pattern transitions, whereas continuous control primarily depends on the evolving object-hand relationship. HINT invokes semantic reasoning only at pattern transitions to resolve the current subtask and target, then maintains this commitment through multi-view grounding and visual tracking. We explore two visual interfaces-image-space semantic highlighting and attention-prior injection-to communicate the tracked intent to the action policy without introducing additional trainable parameters into the foundation action model. Experiments across three long-horizon tasks and out-of-distribution variants show that HINT substantially improves intent understanding, task progress, and end-to-end success across two foundation policies while preserving low-latency control.

机器人操作意图理解长程任务视觉跟踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。