arXiv:2607.19886cs.CV2026-07中稿 · ECCV

融合深度与文本信息,提升热成像转可见光人脸的准确性和一致性。

MTVDiff: Multimodal Conditional Latent Diffusion for Enhanced Thermal-to-Visible Face Translation

论文配图:MTVDiff: Multimodal Conditional Latent Diffusion for Enhanced Thermal-to-Visible Face Translation
图 1 · 摘自论文原文
  • 用双分支注意力融合热成像与深度特征,实现多尺度信息整合。
  • 引入门控文本对齐机制,生成更符合语义的可见光人脸图像。
  • 在MCXFace和SpeakingFaces数据集上性能显著超越现有方法。

热成像转可见光人脸翻译面临几何不连续、语义属性错配和身份退化等挑战。本文提出MTVDiff,一种融合深度与文本信息的多模态潜在扩散框架,以解决上述问题并保留身份特征。核心贡献包括:(1) 双分支交叉注意力融合模块(DBCAF),用于多尺度热成像-深度特征提取与融合;(2) 门控文本到视觉特征对齐机制,实现语义引导生成;(3) 空间特征变换(SFT),支持自适应多模态先验集成。在MCXFace和SpeakingFaces数据集上的大量实验表明,该方法显著优于现有基于GAN和扩散模型的方法,在图像质量和人脸识别性能上均有提升,FID降低最多达48.3%,Rank-1准确率提升最高达8.9%。本工作为复杂光照条件下的人脸识别系统提供了可靠解决方案,推动了跨谱图像翻译的技术前沿。

原文摘要 · Abstract (English)

Thermal-to-visible face translation presents fundamental challenges including geometric discontinuities, semantic attribute mismatches, and identity degradation. We propose MTVDiff, a novel multimodal latent diffusion framework that synergistically integrates depth and textual information to address these limitations while preserving identity characteristics. The MTVDiff framework presents three core technical contributions: (1) a Dual-Branch Cross-Attention Fusion (DBCAF) module for multi-scale thermal-depth feature extraction and fusion; (2) a Gated Text-to-Visual Feature Alignment mechanism for semantically-guided generation; and (3) Spatial Feature Transformations (SFT) for adaptive multimodal prior integration. Extensive experiments on the MCXFace and SpeakingFaces datasets demonstrate that our multimodal approach significantly outperforms existing GAN-based and diffusion-based approaches across multiple metrics, achieving substantial improvements in both image quality and face verification performance, with FID reductions of up to 48.3% and Rank-1 accuracy improvements of up to 8.9\%. Our work provides a robust solution for face recognition systems operating under varying illumination conditions and advances the state-of-the-art in cross-spectral facial image translation through effective multimodal integration.

热成像人脸转换扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。