用大模型生成配对热成像与可见光人脸图,解决数据稀缺问题。
LaPIG: Cross-Modal Generation of Paired Thermal and Visible Facial Images
- 用LLM生成描述词,驱动可见光图合成与热成像转换。
- 在公开数据集上生成的配对图像质量优于现有方法。
- 适合需要多模态人脸数据的研究者使用。
现代机器学习,尤其是人脸识别网络的成功高度依赖高质量、大规模的配对数据集。然而,获取足够数据往往困难且成本高昂。受扩散模型在高质量图像生成和大语言模型(LLMs)进展的启发,我们提出一种名为LaPIG(LLM辅助配对图像生成)的新框架。该框架利用LLM生成的文本描述,构建全面、高质量的可见光与热成像配对图像。方法包含三部分:基于ArcFace嵌入的可见光图像合成、基于潜在扩散模型(LDMs)的热成像转换,以及由LLM生成的图像描述。本方法不仅生成多视角的配对图像以提升数据多样性,还保持了身份信息的一致性。我们在公开数据集上评估该方法,结果表明其性能优于现有方法。
原文摘要 · Abstract (English)
The success of modern machine learning, particularly in facial translation networks, is highly dependent on the availability of high-quality, paired, large-scale datasets. However, acquiring sufficient data is often challenging and costly. Inspired by the recent success of diffusion models in high-quality image synthesis and advancements in Large Language Models (LLMs), we propose a novel framework called LLM-assisted Paired Image Generation (LaPIG). This framework enables the construction of comprehensive, high-quality paired visible and thermal images using captions generated by LLMs. Our method encompasses three parts: visible image synthesis with ArcFace embedding, thermal image translation using Latent Diffusion Models (LDMs), and caption generation with LLMs. Our approach not only generates multi-view paired visible and thermal images to increase data diversity but also produces high-quality paired data while maintaining their identity information. We evaluate our method on public datasets by comparing it with existing methods, demonstrating the superiority of LaPIG.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。