用智能框架提升缺失模态生成质量,让多模态模型更准更可靠。
How Far Are We from Generating Missing Modalities with Foundation Models?
- 构建动态挖掘与自优化框架,精准提取可用模态语义。
- 图像生成FID降低14%以上,文本生成MER降低10%以上。
- 适合需要高精度模态补全的研究者和应用开发者。
多模态基础模型在多种任务中表现卓越,但其在缺失模态重建中的潜力尚未充分探索。本文提出并形式化了三种缺失模态重建范式,对42种模型变体进行了全面评估,涵盖重建精度与下游任务适应性。分析发现,当前模型在细粒度语义提取和生成模态鲁棒验证方面存在明显短板,导致生成结果常出现偏差。为此,我们设计了一种面向缺失模态重建的代理框架,根据输入上下文动态制定模态感知的挖掘策略,以提取更丰富、更具判别性的语义特征;同时引入自精炼机制,通过内部反馈迭代验证并提升生成质量。实验表明,该方法在图像缺失重建上至少降低14%的FID,文本缺失重建上至少降低10%的MER。代码已开源:https://github.com/Guanzhou-Ke/AFM2。
原文摘要 · Abstract (English)
Multimodal foundation models have demonstrated impressive capabilities across diverse tasks. However, their potential as plug-and-play solutions for missing modality reconstruction remains underexplored. To bridge this gap, we identify and formalize three potential paradigms for missing modality reconstruction, and perform a comprehensive evaluation across these paradigms, covering 42 model variants in terms of reconstruction accuracy and adaptability to downstream tasks. Our analysis reveals that current foundation models often fall short in two critical aspects: (i) fine-grained semantic extraction from the available modalities, and (ii) robust validation of generated modalities. These limitations lead to suboptimal and, at times, misaligned generations. To address these challenges, we propose an agentic framework tailored for missing modality reconstruction. This framework dynamically formulates modality-aware mining strategies based on the input context, facilitating the extraction of richer and more discriminative semantic features. In addition, we introduce a self-refinement mechanism, which iteratively verifies and enhances the quality of generated modalities through internal feedback. Experimental results show that our method reduces FID for missing image reconstruction by at least 14\% and MER for missing text reconstruction by at least 10\% compared to baselines. Code are released at: https://github.com/Guanzhou-Ke/AFM2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。