用智能导航器让现成视觉模型自动找好视角,无需标注也能提升新环境表现。
See, Act, Adapt: Active Perception for Unsupervised Cross-Domain Visual Adaptation via Personalized VLM-Guided Agent
- 用视觉语言模型生成动作指令,控制摄像头移动到最佳观察位置。
- 在ReplicaCAD数据集上三项任务性能提升超13%,最高达27.68%。
- 不需重新训练模型或人工标注,适合快速部署到新场景的智能系统。
预训练感知模型在通用图像领域表现优异,但在室内等新环境性能显著下降。传统方法依赖下游数据微调,导致先验知识遗忘且需大量场景标注。本文提出海星(Sea²)范式:不修改感知模块,而是通过智能姿态控制代理优化其使用方式。Sea²保持所有感知模块冻结,训练无需下游标签,仅依靠感知反馈信号引导代理寻找信息丰富的视角。具体地,通过两阶段训练将视觉语言模型(VLM)转化为低层姿态控制器:先在规则探索轨迹上微调,再通过无监督强化学习,利用感知模块输出和置信度构建奖励信号进行策略优化。与以往耦合特定模型或需重训练的数据采集方法不同,Sea²直接复用现成感知模型完成多任务,无需重训练。在三个视觉任务(视觉定位、分割、3D框估计)上验证,在ReplicaCAD数据集上分别取得13.54%、15.92%和27.68%的性能提升。
原文摘要 · Abstract (English)
Pre-trained perception models excel in generic image domains but degrade significantly in novel environments like indoor scenes. The conventional remedy is fine-tuning on downstream data which incurs catastrophic forgetting of prior knowledge and demands costly, scene-specific annotations. We propose a paradigm shift through Sea$^2$ (See, Act, Adapt): rather than adapting the perception modules themselves, we adapt how they are deployed through an intelligent pose-control agent. Sea$^2$ keeps all perception modules frozen, requiring no downstream labels during training, and uses only scalar perceptual feedback to navigate the agent toward informative viewpoints. Specially, we transform a vision-language model (VLM) into a low-level pose controller through a two-stage training pipeline: first fine-tuning it on rule-based exploration trajectories that systematically probe indoor scenes, and then refining the policy via unsupervised reinforcement learning that constructs rewards from the perception module's outputs and confidence. Unlike prior active perception methods that couple exploration with specific models or collect data for retraining them, Sea$^2$ directly leverages off-the-shelf perception models for various tasks without the need for retraining. We conducted experiments on three visual perception tasks, including visual grounding, segmentation and 3D box estimation, with performance improvements of 13.54%, 15.92% and 27.68% respectively on dataset ReplicaCAD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。