arXiv:2605.23281cs.CV2026-05被引 1

通过智能选择最佳深度模型,提升单目深度估计在不同镜头下的表现

DepthAgent: Towards Better Universal Depth Estimation via Sample-wise Expert Selection

论文配图:DepthAgent: Towards Better Universal Depth Estimation via Sample-wise Expert Selection
图 1 · 摘自论文原文
  • 基于视觉-语言代理动态选择适配输入场景的深度模型
  • 在鱼眼和全景图像上显著优于单一模型与固定融合方法
  • 适合需要跨镜头部署深度估计的工程应用

单目度量深度估计在大规模训练和通用相机建模下取得显著进展,但在透视、鱼眼和全景图像等多样相机设置下的鲁棒部署仍具挑战。现有方法通常依赖单一深度模型,忽视不同模型对相机假设的编码差异及其在不同输入域中的最优表现。本文发现深度专家具有强样本级互补性:模型偏好与相机几何高度相关,多模型融合在个体专家不可靠的困难样本上带来最大增益。为此,我们提出 extbf{ ous},一个用于自适应单目深度估计的视觉-语言代理。DepthAgent将现有深度模型视为冻结工具,学习分析场景与相机线索,通过多轮工具调用,为每个输入选择或融合合适专家的预测。为优化此类离散决策以提升稠密几何质量,设计了多奖励强化微调方案,联合鼓励有效工具执行、相机/场景分析、专家选择质量与推理效率。在透视、鱼眼和全景基准上的大量实验表明, ous 始终优于个体专家、固定模型融合与不同选择策略,尤其在困难样本上提升显著,凸显专家选择与融合的关键作用。代码与模型将在发表后公开。

原文摘要 · Abstract (English)

Monocular metric depth estimation has achieved strong progress with large-scale training and universal-camera modeling, yet robust deployment across diverse camera settings, such as perspective, fisheye, and panoramic images, remains challenging. Existing methods typically rely on a single depth estimator, overlooking that different models encode different camera assumptions and perform best under different input domains. In this paper, we show that depth experts exhibit strong sample-wise complementarity: model preference is highly correlated with camera geometry, and multi-model fusion brings the largest gains on difficult samples where individual experts are unreliable. Motivated by these observations, we propose \textbf{\ours}, a vision-language agent for adaptive monocular depth estimation. DepthAgent treats existing depth models as frozen tools and learns to analyze scene and camera cues, invoke suitable experts through multi-turn tool utilization, and select or fuse their predictions for each input. To optimize such discrete decision-making toward dense geometric quality, we design a multi-reward reinforcement fine-tuning scheme that jointly encourages valid tool execution, camera/scene analysis, expert-selection quality, and inference efficiency. Extensive experiments across perspective, fisheye, and panoramic benchmarks show that \ours consistently outperforms individual experts, fixed model fusion, and different selection strategies, with strong improvements on challenging samples, highlighting the critical role of expert selection and fusion. The code and model will be released upon publication.

深度估计视觉语言模型选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。