让机器人根据语言指令自主调整摄像头视角,实现主动感知。
LIME: Learning Intent-aware Camera Motion from Egocentric Video

- 通过自然语言指令预测相机相对位姿,结合视觉与语言理解
- 在真实第一人称视频中挖掘多意图动作监督信号,支持细粒度视角变化
- 可直接从人类录制视频中学习,适用于机器人主动感知任务
自主机器人常需先移动摄像头再执行操作:检查物体、揭示遮挡区域或响应用户意图。尽管视觉-语言导航可将指令转为基础运动,视觉-语言-动作策略可映射至操作行为,但语言驱动的相机运动仍较少被作为独立动作研究。本文提出语言条件下的相机运动生成任务:给定当前RGB观测和自由形式的自然语言意图,预测下一个观测的相对目标相机位姿。该任务具有挑战性:视角变化由潜在感知意图驱动,有效运动可能在不同语义粒度下表现,如进入房间、探看拐角、检查可见物体或揭示遮挡细节。为此,我们从第一人称视频中挖掘多意图相机运动监督数据,将合理意图与观察收益描述配对相对SE(3)目标位姿。提出LIME模型,融合自回归观察收益输出与连续流匹配位姿头,使模型能联合预测下一视图应揭示的内容,并表示多假设目标视角。实验与下游机器人任务表明,LIME能从被动人类视频中学习主动选择相机位姿,将普通第一人称记录转化为意图感知的主动感知监督信号。
原文摘要 · Abstract (English)
Autonomous robots often need to move their camera before they can act: to inspect an object, reveal an occluded region, or obtain a view that responds to a user's intent. While vision-language navigation translates instructions to base motion and vision-language-action policies map instructions to manipulation actions, language-conditioned camera motion remains comparatively underexplored as a first-class action. We formulate language-conditioned camera motion generation: given a current RGB observation and a free-form natural-language intent, predict a relative target camera pose for the next observation. This task is inherently non-trivial: viewpoint changes are driven by latent perceptual intentions, and a valid motion may operate at different semantic granularity, from entering a room to looking around a corner, inspecting a visible object, or revealing an occluded detail. To model this structure, we mine multi-intention camera-motion supervision from egocentric video, pairing plausible intents and observation-gain descriptions with relative SE(3) target poses. We propose LIME, a vision-language camera-motion generator that combines an auto-regressive observation-gain output with a continuous flow-matching pose head. This design lets the model jointly predict what the next view should reveal while representing multi-hypothesis target views. Across experiments and downstream robotic tasks, we show that LIME can learn to actively choose camera poses from passive human video, turning ordinary egocentric recordings into supervision for intent-aware active perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。