arXiv:2602.13588cs.CVcs.AI2026-02

双流协同学习框架,同时提升场景解析与几何视觉任务表现

Two-Stream Interactive Joint Learning of Scene Parsing and Geometric Vision Tasks

  • 双流交互设计:场景流与几何流双向融合特征
  • 无需人工标注对应关系,可利用大规模多视角数据自进化
  • 在三个公开数据集上超越现有最佳方法,适合多任务视觉研究者

受人类视觉系统并行但互动的上下文与空间理解机制启发,本文提出双交互流(TwInS)框架,可同步执行场景解析与几何视觉任务。该框架采用统一通用架构,将场景解析流的多层次上下文特征注入几何流,指导其迭代优化;反向则通过新型跨任务适配器,将解码后的几何特征投影至上下文特征空间,实现选择性异构特征融合,利用丰富的跨视图几何线索增强场景解析。为摆脱对昂贵人工标注对应真值的依赖,TwInS引入定制化半监督训练策略,释放大规模多视角数据潜力,支持无真值条件下持续自我演进。在三个公开数据集上的大量实验验证了核心组件有效性,并证明其性能优于现有最先进方法。源代码将在发表后公开。

原文摘要 · Abstract (English)

Inspired by the human visual system, which operates on two parallel yet interactive streams for contextual and spatial understanding, this article presents Two Interactive Streams (TwInS), a novel bio-inspired joint learning framework capable of simultaneously performing scene parsing and geometric vision tasks. TwInS adopts a unified, general-purpose architecture in which multi-level contextual features from the scene parsing stream are infused into the geometric vision stream to guide its iterative refinement. In the reverse direction, decoded geometric features are projected into the contextual feature space for selective heterogeneous feature fusion via a novel cross-task adapter, which leverages rich cross-view geometric cues to enhance scene parsing. To eliminate the dependence on costly human-annotated correspondence ground truth, TwInS is further equipped with a tailored semi-supervised training strategy, which unleashes the potential of large-scale multi-view data and enables continuous self-evolution without requiring ground-truth correspondences. Extensive experiments conducted on three public datasets validate the effectiveness of TwInS's core components and demonstrate its superior performance over existing state-of-the-art approaches. The source code will be made publicly available upon publication.

场景解析双流网络几何视觉自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。