arXiv:2511.18319cs.AIcs.LG2025-11

用轻量级模型让机器人更准地按指令对齐目标物体

Weakly-supervised Latent Models for Task-specific Visual-Language Control

  • 基于目标状态监督学习隐空间中的动作变化
  • 在无人机对齐任务中达到71%成功率,比直接使用大模型高13个百分点
  • 适合需要精准视觉控制的巡检场景,尤其资源受限时

在危险环境中实现自主巡检,要求智能体能理解高层指令并执行精确控制。关键能力之一是空间定位,例如无人机需将检测到的物体居中于摄像头视野以保证可靠检查。虽然大语言模型提供了自然的目标描述接口,但直接用于视觉控制仅能达到58%的成功率。我们设想,为智能体配备世界模型作为工具,可使其预演候选动作,在空间定位任务中表现更优,但传统世界模型数据和算力需求高。为此,我们提出一种任务特定的隐动态模型,仅通过目标状态监督,在共享隐空间中学习状态相关的动作影响。模型利用全局动作嵌入和互补训练损失,提升学习稳定性。实验表明,该方法在任务中取得71%的成功率,并能泛化到未见图像和指令,证明了紧凑、领域特定的隐动态模型在自主巡检中实现空间对齐的巨大潜力。

原文摘要 · Abstract (English)

Autonomous inspection in hazardous environments requires AI agents that can interpret high-level goals and execute precise control. A key capability for such agents is spatial grounding, for example when a drone must center a detected object in its camera view to enable reliable inspection. While large language models provide a natural interface for specifying goals, using them directly for visual control achieves only 58\% success in this task. We envision that equipping agents with a world model as a tool would allow them to roll out candidate actions and perform better in spatially grounded settings, but conventional world models are data and compute intensive. To address this, we propose a task-specific latent dynamics model that learns state-specific action-induced shifts in a shared latent space using only goal-state supervision. The model leverages global action embeddings and complementary training losses to stabilize learning. In experiments, our approach achieves 71\% success and generalizes to unseen images and instructions, highlighting the potential of compact, domain-specific latent dynamics models for spatial alignment in autonomous inspection.

视觉控制隐空间建模自主巡检

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。