提出新度量方法ID-Sim,更贴近人类对身份的敏感识别能力。
ID-Sim: An Identity-Focused Similarity Metric
- 基于真实与合成数据构建可区分身份的评估指标
- 在跨视角、光照等多变场景中表现优于现有方法
- 适合用于个性化图像生成与身份识别任务的评测
人类对身份具有出色的分辨能力——即使在不同视角或光照条件下,也能轻松区分高度相似的身份。而视觉模型在这一方面表现不足,且由于缺乏面向身份的任务评估指标,个性化图像生成等研究进展受阻。为此,我们提出ID-Sim,一种前馈式度量方法,旨在忠实反映人类对身份的精细敏感性。为构建该指标,我们收集了涵盖多种真实世界场景的高质量图像数据集,并通过生成式合成数据引入可控的细粒度身份与上下文变化。我们在一个全新的统一基准上评估该度量,该基准可同时衡量人类标注的一致性,覆盖身份识别、检索和生成任务。
原文摘要 · Abstract (English)
Humans have remarkable selective sensitivity to identities -- easily distinguishing between highly similar identities, even across significantly different contexts such as diverse viewpoints or lighting. Vision models have struggled to match this capability, and progress toward identity-focused tasks such as personalized image generation is slowed by a lack of identity-focused evaluation metrics. To help facilitate progress, we propose ID-Sim, a feed-forward metric designed to faithfully reflect human selective sensitivity. To build ID-Sim, we curate a high-quality training set of images spanning diverse real-world domains, augmented with generative synthetic data that provides controlled, fine-grained identity and contextual variations. We evaluate our metric on a new unified evaluation benchmark for assessing consistency with human annotations across identity-focused recognition, retrieval, and generative tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。