arXiv:2601.12325cs.CV2026-01被引 1

用轻量超网络提升多模态图像匹配,增强跨域适应性。

Multi-Sensor Matching with HyperNetworks

  • 引入超网络生成通道缩放与偏移参数,实现自适应特征调节。
  • 在VIS-NIR等基准上达到最新水平,且推理成本更低。
  • 适合需要跨模态、跨域鲁棒匹配的研究者使用。

超网络能生成或调节另一网络的权重,为注入上下文与任务条件提供灵活机制,且模型规模增长小。本文利用超网络改进多模态图像块匹配,提出一种轻量级描述符学习架构,通过(i)超网络模块计算自适应的通道缩放与偏移,以及(ii)条件实例归一化,在浅层实现模态特异性适配(如可见光与红外图像)。该方法在保持描述符方法高效推理的同时,显著提升对外观变化的鲁棒性。采用三元组损失与难样本挖掘训练,在VIS-NIR及其他VIS-IR基准上取得当前最优性能,并在多个额外数据集上超越或持平以往高开销方法。为推动领域发展,本文还发布GAP-VIR数据集——一个包含50万对的跨平台(地面/空中)可见光-红外图像块数据集,支持跨域泛化与自适应能力的严格评估。

原文摘要 · Abstract (English)

Hypernetworks are models that generate or modulate the weights of another network. They provide a flexible mechanism for injecting context and task conditioning and have proven broadly useful across diverse applications without significant increases in model size. We leverage hypernetworks to improve multimodal patch matching by introducing a lightweight descriptor-learning architecture that augments a Siamese CNN with (i) hypernetwork modules that compute adaptive, per-channel scaling and shifting and (ii) conditional instance normalization that provides modality-specific adaptation (e.g., visible vs. infrared, VIS-IR) in shallow layers. This combination preserves the efficiency of descriptor-based methods during inference while increasing robustness to appearance shifts. Trained with a triplet loss and hard-negative mining, our approach achieves state-of-the-art results on VIS-NIR and other VIS-IR benchmarks and matches or surpasses prior methods on additional datasets, despite their higher inference cost. To spur progress on domain shift, we also release GAP-VIR, a cross-platform (ground/aerial) VIS-IR patch dataset with 500K pairs, enabling rigorous evaluation of cross-domain generalization and adaptation.

多模态匹配超网络跨域泛化视觉传感器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。