arXiv:2511.10107cs.CV2025-11NeurIPS被引 1

提出RobIA框架,提升立体深度估计在持续变化环境下的适应能力。

RobIA: Robust Instance-aware Continual Test-time Adaptation for Deep Stereo

  • 通过自注意力机制动态路由输入到冻结专家,实现输入感知的轻量适配。
  • 利用伪标签提供密集监督,显著提升在稀疏标注下的泛化性能。
  • 适合需要持续适应真实世界动态变化的立体视觉系统部署。

真实场景中的立体深度估计面临动态域偏移、稀疏或不可靠的监督信号,以及获取密集真值标签成本高昂等挑战。尽管近期测试时自适应(TTA)方法展现出潜力,但多数依赖静态目标域假设和输入无关的适配策略,在持续变化环境下效果受限。本文提出RobIA,一种鲁棒的、实例感知的连续测试时自适应(CTTA)框架,用于立体深度估计。该框架包含两个核心组件:(1) 基于自注意力机制的轻量级多专家模块AttEx-MoE,根据视差几何特性动态分配输入至冻结专家;(2) 基于参数高效微调(PEFT)的稳健自适应教师模型AdaptBN Teacher,通过补充稀疏人工标注生成密集伪监督信号。该策略实现输入特异性灵活性与广泛监督覆盖,显著提升在域偏移下的泛化能力。大量实验表明,RobIA在动态目标域上实现更优的自适应性能,同时保持计算效率。

原文摘要 · Abstract (English)

Stereo Depth Estimation in real-world environments poses significant challenges due to dynamic domain shifts, sparse or unreliable supervision, and the high cost of acquiring dense ground-truth labels. While recent Test-Time Adaptation (TTA) methods offer promising solutions, most rely on static target domain assumptions and input-invariant adaptation strategies, limiting their effectiveness under continual shifts. In this paper, we propose RobIA, a novel Robust, Instance-Aware framework for Continual Test-Time Adaptation (CTTA) in stereo depth estimation. RobIA integrates two key components: (1) Attend-and-Excite Mixture-of-Experts (AttEx-MoE), a parameter-efficient module that dynamically routes input to frozen experts via lightweight self-attention mechanism tailored to epipolar geometry, and (2) Robust AdaptBN Teacher, a PEFT-based teacher model that provides dense pseudo-supervision by complementing sparse handcrafted labels. This strategy enables input-specific flexibility, broad supervision coverage, improving generalization under domain shift. Extensive experiments demonstrate that RobIA achieves superior adaptation performance across dynamic target domains while maintaining computational efficiency.

立体深度测试时自适应持续学习伪标签

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。