arXiv:2606.12303cs.CV2026-06中稿 · ICML被引 3

用1D令牌重构多模态图像融合的全局一致性,提升细节与整体协调性。

From 2D Grids to 1D Tokens: Reforming Shared Representations for Multimodal Image Fusion

论文配图:From 2D Grids to 1D Tokens: Reforming Shared Representations for Multimodal Image Fusion
图 1 · 摘自论文原文
  • 用冻结的预训练图像分词器生成1D令牌,作为全局外观载体
  • 仅稀疏更新关键令牌,实现全局一致性的轻量级控制
  • 在4个基准上同时提升全局一致性和局部保真度,性能最优

多模态图像融合旨在将不同模态的互补信息整合到一个融合图像中,既保留丰富的局部细节,又维持全局一致的外观。现有方法在2D特征网格上构建共享表示,虽擅长建模局部结构,但对图像级全局外观因素的调控能力有限。为此,我们提出一种基于冻结预训练图像分词器的紧凑1D令牌接口,用于建模非局部的外观/基础因素。不同于将分词器作为重建主干,我们的设计将1D令牌空间作为全局载体,同时保留2D空间路径以恢复局部结构。具体提出选择性令牌编辑(STE),仅稀疏更新或替换少量关键令牌,提供轻量级机制以引导全局外观一致性,且不改变融合主干,避免额外损失。在四个常用基准上的实验表明,该方法整体性能最佳,在全局一致性与局部保真度方面均实现持续、多指标提升。

原文摘要 · Abstract (English)

Multimodal image fusion aims to integrate complementary information from different modalities into a fused image that preserves rich local details while maintaining globally consistent appearance. Existing approaches build shared representations on 2D feature grids, which excel at modeling local structures but offer limited leverage over image-level global appearance factors. To balance these objectives, we introduce a compact 1D token interface based on a frozen pretrained image tokenizer for modeling non-local appearance/base factors. Rather than using the tokenizer as a reconstruction backbone, our design uses the 1D token space as a global carrier while retaining the 2D spatial pathway for local structure restoration. Specifically, we introduce Selective Token Editing (STE), which sparsely updates/replaces a small set of critical tokens, providing a lightweight mechanism to steer global appearance coherence while keeping the fusion backbone unchanged and avoiding extra losses. Experiments on four commonly used benchmarks show that our method achieves the best overall performance, with consistent, multi-metric improvements in both global coherence and local fidelity. Project page: https://zju-xyc.github.io/1D-Fusion-Project-Page/

图像融合1D令牌多模态全局一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。