arXiv:2606.15104cs.CV2026-06

用双曲空间融合红外与可见光图像,让语义层次更自然、细节更清晰。

Text-Driven Fusion for Infrared and Visible Images: Achieving Image Scene Adaptation on Hyperbolic Space

论文配图:Text-Driven Fusion for Infrared and Visible Images: Achieving Image Scene Adaptation on Hyperbolic Space
图 1 · 摘自论文原文
  • 用双曲空间建模语义层级,避免传统欧式方法的度量失真。
  • 在IRSTD、FLIR-AI等数据集上优于现有方法,提升显著。
  • 适合需要多模态融合与细粒度语义理解的视觉任务。

红外与可见光图像融合旨在整合互补模态,但现有欧氏方法施加刚性距离度量,扭曲了多模态交互及父子语义层级关系。为此,我们提出一种文本驱动的融合框架,基于双曲流形学习。训练阶段,通过BLIP提取的文本提示作为双曲空间中的拓扑锚点,利用双曲嵌入引导视觉-属性对齐,自然适应不同粒度的语义。借助庞加莱球负曲率带来的指数级体积增长,该方法无缝嵌入层次树结构,编码粗粒度到细粒度语义,而广阔的外围空间防止跨模态融合时的纹理失真。推理阶段,融合过程自动依据输入内容适配,完全无需额外文本输入。实验表明,本方法在基准数据集(IRSTD、FLIR-AI)上超越现有最优方案,代码已开源:https://github.com/Shaoyun2023/TEDFusion。

原文摘要 · Abstract (English)

Infrared and visible image fusion aims to integrate complementary modalities, while existing Euclidean methods impose rigid distance metrics that distort multi-modal interactions and parent-to-child semantic hierarchies. To overcome these limitations, we introduce a text-driven fusion framework empowered by hyperbolic manifold learning. During training, BLIP-extracted text prompts serve as topological anchors within the hyperbolic space, guiding vision-attribute alignment through hyperbolic embeddings that naturally accommodate varying semantic granularities. By exploiting the exponential volume growth dictated by the Poincaré ball's negative curvature, this approach seamlessly embeds hierarchical trees to encode coarse-to-fine semantics without metric saturation, while the vast peripheral space prevents texture distortion during cross-modal fusion. At inference, the fusion process autonomously adapts to input content using the learned text-attribute priors, completely eliminating the need for textual input. Experimental results show our method outperforms state-of-the-art approaches on benchmark datasets, with code available at https://github.com/Shaoyun2023/TEDFusion.

图像融合双曲空间多模态语义层次

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。