arXiv:2608.01794cs.CVcs.AI2026-08中稿 · CVPR

让通用多模态模型具备识别视觉身份的能力,提升生成内容的可信度。

Illuminating Visual Identity in Universal Multimodal Embeddings

论文配图:Illuminating Visual Identity in Universal Multimodal Embeddings
图 1 · 摘自论文原文
  • 设计身份感知采样机制,联合优化通用与身份特征。
  • 在真实与合成数据上构建大规模基准,验证身份识别能力。
  • 适合关注生成内容真实性、跨模态检索的研究者。

通用多模态嵌入(UMEs)旨在将多种模态和任务统一到共享表示空间中。近年来,多模态大语言模型(MLLMs)推动了该领域的显著进展,但视觉身份区分这一关键能力仍被忽视,而它对实例检索、重识别及生成内容的身份保持至关重要。为此,我们提出统一的视觉身份区分(VisID)形式化定义,并引入MVEB(Multimodal Visual Identity Embedding Benchmark),一个涵盖真实世界与合成数据的大规模基准,支持评估与训练。此外,我们设计了一种简单有效的学习框架,通过精心设计的身份感知采样机制,联合优化通用多模态与视觉身份表示。大量实验表明,所提方法成功赋予UMEs强大的身份区分能力,同时保持了优异的通用多模态性能。我们认为这项工作不仅揭示了被忽视的关键能力,也迈向更全面的通用多模态嵌入。

原文摘要 · Abstract (English)

Universal Multimodal Embeddings (UMEs) aim to unify various modalities and tasks into a shared representation space. In recent years, this field has witnessed substantial progress driven by the development of Multimodal Large Language Models (MLLMs). However, a crucial capability, visual identity discrimination, remains underexplored in existing UME methods, despite its critical role in a wide range of tasks, including instance retrieval, re-identification, and identity preservation in AI-generated content. To bridge this gap, we propose a unified formulation for visual identity discrimination~(VisID) and introduce $\textbf{MVEB}$ ($\textbf{M}$ultimodal $\textbf{V}$isual Identity $\textbf{E}$mbedding $\textbf{B}$enchmark), a large-scale benchmark curated from both real-world and synthetic datasets to support evaluation and training. Furthermore, we present a simple yet effective learning framework that jointly optimizes general multimodal and visual identity representations through a carefully designed identity-aware sampling mechanism. Extensive experiments demonstrate that our approach successfully endows UMEs with strong identity discrimination capability and maintains competitive general multimodal performance. We believe this work not only illuminates a critical yet neglected capability, but also takes a step toward more holistic universal multimodal embeddings. Code and data are available at \href{https://chrisclear3.github.io/MVEB}{MVEB}.

多模态身份识别生成内容嵌入学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。