提出可控后验桥学习框架,提升多任务密集预测的可靠性。
$\mathcal{B}^{3}$-Net: Controlled Posterior Bridge Learning for Multi-Task Dense Prediction

- 分三步构建可信证据桥:估精度、加权融合、有限重分配
- 在多个数据集上优于主流方法,显著减少负迁移
- 适合需要高可靠多任务推理的视觉系统设计者
多任务密集预测在统一模型中处理语义分割、深度估计、法向量估计和边缘检测等像素级任务。现有解码器交互方法通过注意力、提示、路由、扩散、Mamba或桥接特征交换任务信息,但大多隐式组织证据。它们通常按相似性或亲和度融合特征,未显式建模证据可靠性在任务与空间位置间的差异,导致不可靠证据污染共享表示,加剧负迁移。本文提出$$\mathcal{B}^{3}$-Net$,一种受控后验桥学习框架。方法将解码器交互分解为:可靠性估计、后验桥构建和有界重分配。精度场估计器通过任务对齐与局部变化估算块级证据精度;后验桥操作器通过异方差证据融合构建精度加权后验桥,生成比均匀或启发式混合更可靠的共享状态;收缩调度操作器通过有界更新将桥分配至各任务分支,减少不受控特征注入。在NYUD-v2、PASCAL-Context和Cityscapes上的实验表明,$$\mathcal{B}^{3}$-Net$在代表性CNN、Transformer、扩散、Mamba及桥特征方法间达到竞争性或更优的权衡。骨干匹配对比与详尽分析进一步验证,性能提升源于受控后验桥学习,而非骨干能力或解码器规模。
原文摘要 · Abstract (English)
Multi-task dense prediction solves complementary pixel-level tasks in a unified model, such as semantic segmentation, depth estimation, surface normal estimation, and edge detection. Existing decoder-side interactions use attention, prompts, routing, diffusion, Mamba, or bridge features to exchange task evidence, but most of them organize this evidence implicitly. They usually fuse task features by similarity or affinity, without explicitly modeling that evidence reliability varies across tasks and spatial locations. As a result, unreliable evidence may contaminate the shared representation and intensify negative transfer. We propose $\mathcal{B}^{3}$-Net, a controlled posterior bridge learning framework for multi-task dense prediction. Our method decomposes decoder-side interaction into reliability estimation, posterior bridge construction, and bounded redistribution. The Precision Field Estimator estimates patch-wise evidence precision from task-reference alignment and local variation. The Posterior Bridge Operator builds a precision-weighted posterior bridge through heteroscedastic evidence fusion, yielding a shared state more reliable than uniform or heuristic mixtures. The Contractive Dispatch Operator redistributes the bridge to each task branch through a bounded update, reducing uncontrolled feature injection. Experiments on NYUD-v2, PASCAL-Context, and Cityscapes show that $\mathcal{B}^{3}$-Net achieves competitive or superior trade-offs over representative CNN-, Transformer-, diffusion-, Mamba-, and bridge-feature-based methods. Backbone-matched comparisons and extensive analyses further verify that the gains arise from controlled posterior bridge learning rather than backbone capacity or decoder scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。