让机器人根据语言指令主动选视角,看清任务所需信息。
I-Perceive: A Foundation Model for Active Perception with Language Instructions
- 融合视觉语言与几何推理,实现语言控制的主动感知
- 在真实与仿真数据上训练,预测准确率显著超越基线
- 零样本泛化强,可闭环迭代优化视角,适合复杂场景应用
主动感知——机器人主动选择视角以获取任务相关的信息——是实现场景中稳健运行的关键。然而,现有方法通常局限于固定目标或受限环境,难以泛化到自然语言描述的开放性感知意图。我们提出 I-Perceive,一个面向大规模室内环境的语言条件主动感知基础模型。给定查询图像、一组上下文图像和自然语言指令,I-Perceive 预测一个满足指定感知意图的 6D 相机位姿。该模型结合视觉-语言路径进行语义定位,以及几何推理路径实现多视角三维理解,并通过多层语义融合连接,支持语言条件下的几何推理。为支持规模化训练,我们利用自动化流程,从真实场景扫描数据与仿真环境构建大规模语言-视角对数据集。大量实验表明,I-Perceive 在预测精度、视角可行性及指令对齐方面显著优于强基线。模型展现出强大的零样本泛化能力,可实现闭环主动感知,通过连续交互逐步优化视角。
原文摘要 · Abstract (English)
Active perception - the ability of a robot to proactively select viewpoints to acquire task-relevant information - is essential for robust operation in real-world environments. However, existing approaches are typically limited to fixed objectives or constrained settings, and struggle to generalize to open-ended perception intents specified in natural language. We propose I-Perceive, a foundation model for language-conditioned active perception in large-scale indoor environments. Given a query image, a set of context images, and a natural language instruction, I-Perceive predicts a 6D camera pose that fulfills the specified perception intent. The model integrates a vision-language pathway for semantic grounding with a geometric reasoning pathway for multi-view 3D understanding, connected via multi-layer semantic fusion to enable language-conditioned geometric reasoning. To support scalable training, we construct a large-scale dataset of language-viewpoint pairs from both real-world scene-scanning data and simulated environments using an automated pipeline. Extensive experiments demonstrate that I-Perceive significantly outperforms strong baselines on prediction accuracy, viewpoint feasibility, and instructions alignment. The model exhibits strong zero-shot generalization to unseen scenes and instructions, and enables closed-loop active perception, progressively refining viewpoints over sequential interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。