arXiv:2606.03470cs.CV2026-06

让人脸和发型跨模态匹配,实现精准形象检索

Mixed-Modality Dual Face-Hair Retrieval

论文配图:Mixed-Modality Dual Face-Hair Retrieval
图 1 · 摘自论文原文
  • 用图像或文字指定发型,与人脸联合检索
  • 构建超18万条标注样本的首个跨模态发型检索基准
  • 通过嵌入分解与融合,实现身份与发型解耦控制

我们提出双参考跨模态人脸-发型检索(DFHR),一种新型图像检索任务:查询由一张人脸图像(表征身份)和一张发型参考图像或文本组成。不同于以往检索设置,DFHR需在语义独立的身份与发型之间进行跨组件推理,且源自异构模态。该任务要求局部特征解耦、跨模态语义对齐以及统一嵌入空间内的混合模态组合。我们构建了首个基准数据集DFHR-Bench,包含超过18万条标注三元组,覆盖双图像与图像-文本两种设置,通过多阶段标注协议确保语义与身份完整性。我们进一步提出MFHC(多模态人脸-发型组合器),通过标记注入与多视角监督融合解耦的身份与发型嵌入。DFHR与DFHR-Bench共同建立了一种跨模态、身份感知、属性可控的视觉检索新范式。

原文摘要 · Abstract (English)

We introduce Dual Face-Hair Retrieval (DFHR), a new mixed-modality dual-reference task in image retrieval where a query consists of a face image specifying identity and a hairstyle reference expressed as either an image or text. Unlike prior retrieval settings, DFHR requires cross-component reasoning between two semantically independent attributes -- identity and hairstyle -- originating from heterogeneous modalities. This formulation demands localized feature disentanglement, cross-modal semantic alignment, and mixed-modality composition within a unified embedding space. We construct DFHR-Bench, the first benchmark for mixed-modality face-hair retrieval, comprising over 180K annotated triplets across dual-image and image-text settings, built via a multi-stage annotation protocol ensuring semantic and identity integrity. We further propose MFHC (Multimodal Face-Hair Combiner), a unified framework that fuses disentangled identity and hairstyle embeddings through token injection and multi-view supervision. DFHR and DFHR-Bench together establish a new paradigm for identity-aware, attribute-controllable visual retrieval across modalities.

跨模态检索人脸生成属性控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。